Instructions to use Ninnix96/Qwengram-0.8B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Ninnix96/Qwengram-0.8B with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Ninnix96/Qwengram-0.8B:Q4_K_M # Run inference directly in the terminal: llama cli -hf Ninnix96/Qwengram-0.8B:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Ninnix96/Qwengram-0.8B:Q4_K_M # Run inference directly in the terminal: llama cli -hf Ninnix96/Qwengram-0.8B:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Ninnix96/Qwengram-0.8B:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf Ninnix96/Qwengram-0.8B:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Ninnix96/Qwengram-0.8B:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf Ninnix96/Qwengram-0.8B:Q4_K_M
Use Docker
docker model run hf.co/Ninnix96/Qwengram-0.8B:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use Ninnix96/Qwengram-0.8B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Ninnix96/Qwengram-0.8B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Ninnix96/Qwengram-0.8B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Ninnix96/Qwengram-0.8B:Q4_K_M
- Ollama
How to use Ninnix96/Qwengram-0.8B with Ollama:
ollama run hf.co/Ninnix96/Qwengram-0.8B:Q4_K_M
- Unsloth Desktop
- Pi
How to use Ninnix96/Qwengram-0.8B with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Ninnix96/Qwengram-0.8B:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Ninnix96/Qwengram-0.8B:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use Ninnix96/Qwengram-0.8B with Docker Model Runner:
docker model run hf.co/Ninnix96/Qwengram-0.8B:Q4_K_M
- Lemonade
How to use Ninnix96/Qwengram-0.8B with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Ninnix96/Qwengram-0.8B:Q4_K_M
Run and chat with the model
lemonade run user.Qwengram-0.8B-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use Ninnix96/Qwengram-0.8B with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Ninnix96/Qwengram-0.8B:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Ninnix96/Qwengram-0.8B:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Ninnix96/Qwengram-0.8B with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Ninnix96/Qwengram-0.8B:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Ninnix96/Qwengram-0.8B:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Qwengram-0.8B
Frozen Qwen3.5-0.8B plus an R=1 reader at decoder IDX2/IDX8 and linear750 dynamic arbitration, using external Qwen3.8-Flash-Next PLE memory.
Canonical balanced endpoint: REAL-15M + linear750 (15,000,064 reader tokens; 749,568 calibration tokens). The matched milestone study found lower full-validation, five-domain mean and LAMBADA NLL at 15M than at 10M, with paired 95% intervals below zero. Benchmark accuracy differences were not established.
The canonical checkpoint reduces frozen full-validation perplexity by 5.048% (18.2759 stock to 17.3534). The 20M research endpoint improves aggregate NLL further but regresses on math, so 15M remains the balanced choice. Its evaluation and checkpoint provenance are preserved under research-20m/. See the milestone comparisons and paired intervals.
GGUF files
| Precision | Quantization | Download | Size |
|---|---|---|---|
| 4-bit | Q4_K_M | QwenGram-0.8B-Q4_K_M.gguf | 584 MB |
| 6-bit | Q6_K | QwenGram-0.8B-Q6_K.gguf | 688 MB |
| 8-bit | Q8_0 | QwenGram-0.8B-Q8_0.gguf | 876 MB |
| 16-bit | BF16 | QwenGram-0.8B-BF16.gguf | 1.60 GB |
Sizes use decimal MB/GB and exclude the required PLE sidecar. Every GGUF contains the text backbone and 11 FP32 reader/arbiter tensors. The reader and arbiter stay FP32 in every precision. The PLE is a required external file, not embedded in these GGUFs. SHA256.json records artifact sizes and hashes.
The original combined BF16 export and its configuration are in
safetensors/. The canonical reader and arbiter are also
available separately as reader.safetensors and arbiter.pt.
Required PLE sidecar
Download Ivan Fioravanti's Q4_1 PLE GGUF.
Credit for this PLE conversion belongs to Ivan. Its SHA256 is
66db3ab390f4dd5063ecc89cc180f4713898577682347001bf64ab8e328527a1.
The approximately 32 GB file is mapped on the host; only selected rows are
dequantized for each token. It is not loaded as a 32 GB GPU allocation.
Build and run
Use the Qwengram llama.cpp fork,
commit 068fcb42662453bec15298ba6bd59f552190a468, which supports both 0.8B and 2B:
git clone https://github.com/Ninnix/llama.cpp-qwengram.git
cd llama.cpp-qwengram
git checkout 068fcb42662453bec15298ba6bd59f552190a468
cmake -S . -B build-qwengram-cpu -DCMAKE_BUILD_TYPE=Release -DLLAMA_BUILD_EXAMPLES=ON
cmake --build build-qwengram-cpu -j --target llama-completion
export QWENGRAM_PLE=/path/to/Qwen3.8-Flash-Next-PLE-Q4_1.gguf
build-qwengram-cpu/bin/llama-completion -m /path/to/QwenGram-0.8B-Q8_0.gguf -p 'The capital of France is' -n 16 -no-cnv -ngl 0
For Vulkan, build with -DGGML_VULKAN=ON. On the tested AMD BC-250, BF16 and Q8_0
with -ngl 99 matched the CPU first-token check. Q6_K matched the CPU eight-token
greedy continuation with -ngl 99. Q4_K_M produced a different continuation
with full offload; use -ngl 25 on that device. These short generation checks do
not establish broad GPU parity.
Stock upstream llama.cpp and unmodified Transformers do not execute the custom reader. The fork implements reader injection, arbitration and deterministic n-gram addressing. MTP and embedding-only inputs are unsupported for this Qwengram runtime. See the runtime guide.
Frozen evaluation
Canonical REAL-15M + linear750, using the original FP8 PLE and frozen Kaggle evaluation suite.
| Metric | Canonical Qwengram-0.8B |
|---|---|
| Full-validation NLL | 2.853786 |
| Full-validation perplexity | 17.3534 |
| General NLL | 3.134094 |
| Code NLL | 1.495201 |
| Math NLL | 1.402346 |
| Scientific NLL | 2.244491 |
| Multilingual NLL | 3.647268 |
| Five-domain mean NLL | 2.384680 |
| LAMBADA-1000 NLL | 2.217277 |
| LAMBADA-1000 accuracy | 45.6% |
| HellaSwag-1000 accuracy | 40.0% |
Frozen stock full-validation NLL was 2.905585 (perplexity 18.2759). These frozen study metrics are not quantized GGUF measurements. Per-block and per-example results, frozen evaluation identities and paired milestone comparisons are in evaluation/.
GGUF runtime retention
The matched CPU test scores 8,128 tokens from the first 64 consecutive 256-token
WikiText-2 raw test chunks. All runs use eight threads and context/batch/microbatch
256. Reader gain is NLL(stock) - NLL(Qwengram); retention divides each quantized
gain by the BF16 gain. Paired 95% intervals use 10,000 resamples of 16 consecutive
four-chunk blocks, seed 1234. The external PLE is Ivan Fioravanti's Q4_1 sidecar.
| Precision | Stock NLL | Qwengram NLL | Reader gain [95% CI] | Gain retention [95% CI] | Perplexity reduction vs stock |
|---|---|---|---|---|---|
| BF16 | 2.890026 | 2.822039 | 0.067987 [0.058500, 0.077155] | 100% | 6.57% |
| Q8_0 | 2.890255 | 2.822870 | 0.067385 [0.057686, 0.076578] | 99.1% [96.5%, 101.9%] | 6.52% |
| Q6_K | 2.900546 | 2.831509 | 0.069037 [0.059437, 0.078186] | 101.5% [98.3%, 104.7%] | 6.67% |
| Q4_K_M | 2.936757 | 2.874906 | 0.061851 [0.053637, 0.069662] | 91.0% [85.6%, 96.2%] | 6.00% |
Within each stock/Qwengram pair, all 335 backbone tensors and nine tokenizer
fields match exactly; all 11 reader and arbiter tensors remain bit-exact FP32.
The original BF16/Q8_0/Q4_K_M runtime runs used fork commit
1c5053f23e99341bee106c1232af90f57ed2cfdd; Q6_K was quantized and tested with
068fcb42662453bec15298ba6bd59f552190a468. See Q6 verification
and generation outputs.
See the matched stock-versus-Qwengram runtime report for hashes, commands, logs and per-chunk results. It uses a fixed WikiText-2 slice and a quantized Q4_1 sidecar; the Kaggle scores above use the original FP8 PLE and the frozen study suite. They are separate benchmarks.
Provenance
| Artifact | SHA-256 |
|---|---|
| REAL-15M reader weights | e4a760163ec07568178ab48aa533235a9af878183caf120d4e29ce1c8ce4b9dc |
| REAL-15M linear750 gate | 54b7a98dd5a6b0fcde69deddc8ad2efabfe86daaeca840e949c74ed3bd6d8138 |
The target model revision is 2fc06364715b967f1860aea9cf38778875588b17.
The study PLE is Qwen/Qwen3.8-Flash-Next-FP8, revision
236dfdf285828023ca3bcd3f37366c58a3469b13 (about 48.7 GiB).
Source reader dataset: ninnix/qwen-ple-reader-15m-milestone. Source gate dataset:
ninnix/qwen-ple-reader-scale-15m-arbitration-results.
Full provenance is in qwengram-0.8b.json, the evaluation
artifacts, and runtime/. The research code and results
describe the training study. The original base model and its license are from
Qwen. This is an experimental text-generation release; vision has not been validated.
- Downloads last month
- 3,522
4-bit
6-bit
8-bit
16-bit