Kev-27B NF4

This is a self-contained bitsandbytes NF4, double-quantized fork of Kev-27B. It includes the quantized backbone, original Kev LoRA adapter, and calibrated pointer head in this one repository. It answers typed decision questions through Kev's /v1/systemone API. It is not a chat or text-generation model, and NF4 is not NVIDIA NVFP4.

The repository contains only the quantized backbone, Kev adapter and head, tokenizer, benchmark reports, and the small runtime patch needed to load the quantized backbone. It does not contain BF16 backbone weights. The Qwen base is pinned to revision 1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0; the Kev source used for the measurements is commit 3e1cd3bb588a388a06827443380befece23e68c7.

Measured quality

All rows were scored through the live /v1/systemone service with the upstream python -m kev.benchmark --remote runner, two requests in flight, on the complete development partitions of Kev's frozen suites. No records were rejected or truncated. The 16-bit reference figures are from the original Kev-27B release evidence and model card. That evaluation used its own path, so this is not a paired served-BF16-versus-served-NF4 experiment.

Measure NF4 Kev-27B served Published 16-bit Kev-27B Published Kev-9B
transfer-v4 development clean accuracy (656 questions) 0.858 0.848 0.822
transfer-v4 development clean Brier 0.224 0.229 0.264
transfer-v4 development clean ECE 0.041 0.043 0.042
transfer-v4 development coverage at ≤5% error 0.601 0.543 0.447
transfer-v9 development MMLU-Pro accuracy (200 questions) 0.625 0.665 0.515
transfer-v9 development unknowable items answered at ≥0.9 0.000 0.000 0.000
decision-v7 development clean accuracy (1,264 questions) 0.866 0.866 0.872
decision-v7 development clean Brier 0.184 0.185 0.184
decision-v7 development clean ECE 0.014 0.019 0.032
decision-v7 development coverage at ≤5% error 0.777 0.780 0.777

Known quality tradeoffs: MMLU-Pro is four percentage points below the published 16-bit Kev-27B development result and below Kev's documented 0.65 model-selection criterion. On 200 questions, that is eight additional wrong answers. Decision-v7 development accuracy is 0.6 percentage points below published Kev-9B. NF4 improves on Kev-9B by 3.7 points on transfer-v4 development accuracy and 11 points on MMLU-Pro, with lower transfer-v4 Brier. These are comparisons to published evaluations, not paired runs on the same serving hardware. The earlier 30-record API smoke comparison found no answer flips but a maximum probability change of 0.103 against this deployment's BF16 service, so probability parity is not guaranteed.

Full upstream benchmark reports are in benchmarks/; each report includes the suite hash, task breakdown, calibration, selective-risk metrics, and API coverage.

The saved package was loaded into a second Kev service and checked against the running NF4 service on 30 records and 40 answers: zero answer flips and zero probability difference in the returned API values.

Spark serving measurements

On one NVIDIA GB10 Spark, the live NF4 service allocated 17,818 MiB GPU memory when healthy and 17,888 MiB after all suites, versus 63,564 MiB measured for the local BF16 service when idle. The NF4 numbers include the pointer head, adapter, and serving buffers, not just packed weight files. The container used around 23 GiB host RAM after loading; Linux file cache is reclaimable.

Suite Records Wall time Records/s Median API latency
transfer-v4 development 764 311 s 2.46 743 ms
transfer-v9 development 1,264 582 s 2.17 820 ms
decision-v7 development 1,204 760 s 1.58 928 ms

These are end-to-end benchmark measurements with concurrency 2 on GB10; the published H200 model-time figures use different hardware and request shapes.

Run

The bundled adapter/ metadata points to the bundled base/. Download the whole repository and run from its root. This package was verified with Python, CUDA, transformers==5.17.0, bitsandbytes==0.50.2, and peft==0.21.0 on a GB10. Kev must be checked out at the pinned commit because patch_kev.py uses exact source matches.

git clone https://github.com/jaredpalmer/kev.git
git -C kev checkout 3e1cd3bb588a388a06827443380befece23e68c7
python -m pip install -e './kev[serve]' bitsandbytes==0.50.2
python -m pip install 'transformers==5.17.0' 'peft==0.21.0'
hf download TheCulliganMan/kev-27b-nf4 --local-dir kev-27b-nf4
python kev-27b-nf4/patch_kev.py
cd kev-27b-nf4
./serve.sh

The service listens on port 8428 and keeps Kev's kev-latest API model name. serve.sh disables fused weights and CUDA graphs for this NF4 path. Do not merge or cast the adapter into the quantized backbone.

Sources and license

The base is Qwen/Qwen3.8-27B; the decision adapter and head are from jaredpalmer/kev-27b. Both upstream weight releases are Apache-2.0. The evaluation suites and benchmark runner come from the Kev repository. Attribution and source links are provided for the upstream work.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for TheCulliganMan/kev-27b-nf4

Base model

Qwen/Qwen3.8-27B
Finetuned
(458)
this model