Qwengram-0.8B

Qwengram-0.8B logo

Frozen Qwen3.5-0.8B plus an R=1 reader at decoder IDX2/IDX8 and linear750 dynamic arbitration, using external Qwen3.8-Flash-Next PLE memory.

Canonical balanced endpoint: REAL-15M + linear750 (15,000,064 reader tokens; 749,568 calibration tokens). The matched milestone study found lower full-validation, five-domain mean and LAMBADA NLL at 15M than at 10M, with paired 95% intervals below zero. Benchmark accuracy differences were not established.

The canonical checkpoint reduces frozen full-validation perplexity by 5.048% (18.2759 stock to 17.3534). The 20M research endpoint improves aggregate NLL further but regresses on math, so 15M remains the balanced choice. Its evaluation and checkpoint provenance are preserved under research-20m/. See the milestone comparisons and paired intervals.

GGUF files

Precision Quantization Download Size
4-bit Q4_K_M QwenGram-0.8B-Q4_K_M.gguf 584 MB
6-bit Q6_K QwenGram-0.8B-Q6_K.gguf 688 MB
8-bit Q8_0 QwenGram-0.8B-Q8_0.gguf 876 MB
16-bit BF16 QwenGram-0.8B-BF16.gguf 1.60 GB

Sizes use decimal MB/GB and exclude the required PLE sidecar. Every GGUF contains the text backbone and 11 FP32 reader/arbiter tensors. The reader and arbiter stay FP32 in every precision. The PLE is a required external file, not embedded in these GGUFs. SHA256.json records artifact sizes and hashes.

The original combined BF16 export and its configuration are in safetensors/. The canonical reader and arbiter are also available separately as reader.safetensors and arbiter.pt.

Required PLE sidecar

Download Ivan Fioravanti's Q4_1 PLE GGUF. Credit for this PLE conversion belongs to Ivan. Its SHA256 is 66db3ab390f4dd5063ecc89cc180f4713898577682347001bf64ab8e328527a1. The approximately 32 GB file is mapped on the host; only selected rows are dequantized for each token. It is not loaded as a 32 GB GPU allocation.

Build and run

Use the Qwengram llama.cpp fork, commit 068fcb42662453bec15298ba6bd59f552190a468, which supports both 0.8B and 2B:

git clone https://github.com/Ninnix/llama.cpp-qwengram.git
cd llama.cpp-qwengram
git checkout 068fcb42662453bec15298ba6bd59f552190a468
cmake -S . -B build-qwengram-cpu -DCMAKE_BUILD_TYPE=Release -DLLAMA_BUILD_EXAMPLES=ON
cmake --build build-qwengram-cpu -j --target llama-completion
export QWENGRAM_PLE=/path/to/Qwen3.8-Flash-Next-PLE-Q4_1.gguf
build-qwengram-cpu/bin/llama-completion -m /path/to/QwenGram-0.8B-Q8_0.gguf -p 'The capital of France is' -n 16 -no-cnv -ngl 0

For Vulkan, build with -DGGML_VULKAN=ON. On the tested AMD BC-250, BF16 and Q8_0 with -ngl 99 matched the CPU first-token check. Q6_K matched the CPU eight-token greedy continuation with -ngl 99. Q4_K_M produced a different continuation with full offload; use -ngl 25 on that device. These short generation checks do not establish broad GPU parity.

Stock upstream llama.cpp and unmodified Transformers do not execute the custom reader. The fork implements reader injection, arbitration and deterministic n-gram addressing. MTP and embedding-only inputs are unsupported for this Qwengram runtime. See the runtime guide.

Frozen evaluation

Canonical REAL-15M + linear750, using the original FP8 PLE and frozen Kaggle evaluation suite.

Metric Canonical Qwengram-0.8B
Full-validation NLL 2.853786
Full-validation perplexity 17.3534
General NLL 3.134094
Code NLL 1.495201
Math NLL 1.402346
Scientific NLL 2.244491
Multilingual NLL 3.647268
Five-domain mean NLL 2.384680
LAMBADA-1000 NLL 2.217277
LAMBADA-1000 accuracy 45.6%
HellaSwag-1000 accuracy 40.0%

Frozen stock full-validation NLL was 2.905585 (perplexity 18.2759). These frozen study metrics are not quantized GGUF measurements. Per-block and per-example results, frozen evaluation identities and paired milestone comparisons are in evaluation/.

GGUF runtime retention

The matched CPU test scores 8,128 tokens from the first 64 consecutive 256-token WikiText-2 raw test chunks. All runs use eight threads and context/batch/microbatch 256. Reader gain is NLL(stock) - NLL(Qwengram); retention divides each quantized gain by the BF16 gain. Paired 95% intervals use 10,000 resamples of 16 consecutive four-chunk blocks, seed 1234. The external PLE is Ivan Fioravanti's Q4_1 sidecar.

Precision Stock NLL Qwengram NLL Reader gain [95% CI] Gain retention [95% CI] Perplexity reduction vs stock
BF16 2.890026 2.822039 0.067987 [0.058500, 0.077155] 100% 6.57%
Q8_0 2.890255 2.822870 0.067385 [0.057686, 0.076578] 99.1% [96.5%, 101.9%] 6.52%
Q6_K 2.900546 2.831509 0.069037 [0.059437, 0.078186] 101.5% [98.3%, 104.7%] 6.67%
Q4_K_M 2.936757 2.874906 0.061851 [0.053637, 0.069662] 91.0% [85.6%, 96.2%] 6.00%

Within each stock/Qwengram pair, all 335 backbone tensors and nine tokenizer fields match exactly; all 11 reader and arbiter tensors remain bit-exact FP32. The original BF16/Q8_0/Q4_K_M runtime runs used fork commit 1c5053f23e99341bee106c1232af90f57ed2cfdd; Q6_K was quantized and tested with 068fcb42662453bec15298ba6bd59f552190a468. See Q6 verification and generation outputs.

See the matched stock-versus-Qwengram runtime report for hashes, commands, logs and per-chunk results. It uses a fixed WikiText-2 slice and a quantized Q4_1 sidecar; the Kaggle scores above use the original FP8 PLE and the frozen study suite. They are separate benchmarks.

Provenance

Artifact SHA-256
REAL-15M reader weights e4a760163ec07568178ab48aa533235a9af878183caf120d4e29ce1c8ce4b9dc
REAL-15M linear750 gate 54b7a98dd5a6b0fcde69deddc8ad2efabfe86daaeca840e949c74ed3bd6d8138

The target model revision is 2fc06364715b967f1860aea9cf38778875588b17. The study PLE is Qwen/Qwen3.8-Flash-Next-FP8, revision 236dfdf285828023ca3bcd3f37366c58a3469b13 (about 48.7 GiB). Source reader dataset: ninnix/qwen-ple-reader-15m-milestone. Source gate dataset: ninnix/qwen-ple-reader-scale-15m-arbitration-results.

Full provenance is in qwengram-0.8b.json, the evaluation artifacts, and runtime/. The research code and results describe the training study. The original base model and its license are from Qwen. This is an experimental text-generation release; vision has not been validated.

Downloads last month
3,522
GGUF
Model size
0.8B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

4-bit

6-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Ninnix96/Qwengram-0.8B

Quantized
(305)
this model

Space using Ninnix96/Qwengram-0.8B 1

Collection including Ninnix96/Qwengram-0.8B