Kev-27B NF4
This is a self-contained bitsandbytes NF4, double-quantized fork of Kev-27B. It includes the quantized backbone, original Kev LoRA adapter, and calibrated pointer head in this one repository. It answers typed decision questions through Kev's /v1/systemone API. It is not a chat or text-generation model, and NF4 is not NVIDIA NVFP4.
The repository contains only the quantized backbone, Kev adapter and head, tokenizer, benchmark reports, and the small runtime patch needed to load the quantized backbone. It does not contain BF16 backbone weights. The Qwen base is pinned to revision 1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0; the Kev source used for the measurements is commit 3e1cd3bb588a388a06827443380befece23e68c7.
Measured quality
All rows were scored through the live /v1/systemone service with the upstream python -m kev.benchmark --remote runner, two requests in flight, on the complete development partitions of Kev's frozen suites. No records were rejected or truncated. The 16-bit reference figures are from the original Kev-27B release evidence and model card. That evaluation used its own path, so this is not a paired served-BF16-versus-served-NF4 experiment.
| Measure | NF4 Kev-27B served | Published 16-bit Kev-27B | Published Kev-9B |
|---|---|---|---|
| transfer-v4 development clean accuracy (656 questions) | 0.858 | 0.848 | 0.822 |
| transfer-v4 development clean Brier | 0.224 | 0.229 | 0.264 |
| transfer-v4 development clean ECE | 0.041 | 0.043 | 0.042 |
| transfer-v4 development coverage at ≤5% error | 0.601 | 0.543 | 0.447 |
| transfer-v9 development MMLU-Pro accuracy (200 questions) | 0.625 | 0.665 | 0.515 |
| transfer-v9 development unknowable items answered at ≥0.9 | 0.000 | 0.000 | 0.000 |
| decision-v7 development clean accuracy (1,264 questions) | 0.866 | 0.866 | 0.872 |
| decision-v7 development clean Brier | 0.184 | 0.185 | 0.184 |
| decision-v7 development clean ECE | 0.014 | 0.019 | 0.032 |
| decision-v7 development coverage at ≤5% error | 0.777 | 0.780 | 0.777 |
Known quality tradeoffs: MMLU-Pro is four percentage points below the published 16-bit Kev-27B development result and below Kev's documented 0.65 model-selection criterion. On 200 questions, that is eight additional wrong answers. Decision-v7 development accuracy is 0.6 percentage points below published Kev-9B. NF4 improves on Kev-9B by 3.7 points on transfer-v4 development accuracy and 11 points on MMLU-Pro, with lower transfer-v4 Brier. These are comparisons to published evaluations, not paired runs on the same serving hardware. The earlier 30-record API smoke comparison found no answer flips but a maximum probability change of 0.103 against this deployment's BF16 service, so probability parity is not guaranteed.
Full upstream benchmark reports are in benchmarks/; each report includes the suite hash, task breakdown, calibration, selective-risk metrics, and API coverage.
The saved package was loaded into a second Kev service and checked against the running NF4 service on 30 records and 40 answers: zero answer flips and zero probability difference in the returned API values.
Spark serving measurements
On one NVIDIA GB10 Spark, the live NF4 service allocated 17,818 MiB GPU memory when healthy and 17,888 MiB after all suites, versus 63,564 MiB measured for the local BF16 service when idle. The NF4 numbers include the pointer head, adapter, and serving buffers, not just packed weight files. The container used around 23 GiB host RAM after loading; Linux file cache is reclaimable.
| Suite | Records | Wall time | Records/s | Median API latency |
|---|---|---|---|---|
| transfer-v4 development | 764 | 311 s | 2.46 | 743 ms |
| transfer-v9 development | 1,264 | 582 s | 2.17 | 820 ms |
| decision-v7 development | 1,204 | 760 s | 1.58 | 928 ms |
These are end-to-end benchmark measurements with concurrency 2 on GB10; the published H200 model-time figures use different hardware and request shapes.
Run
The bundled adapter/ metadata points to the bundled base/. Download the whole repository and run from its root. This package was verified with Python, CUDA, transformers==5.17.0, bitsandbytes==0.50.2, and peft==0.21.0 on a GB10. Kev must be checked out at the pinned commit because patch_kev.py uses exact source matches.
git clone https://github.com/jaredpalmer/kev.git
git -C kev checkout 3e1cd3bb588a388a06827443380befece23e68c7
python -m pip install -e './kev[serve]' bitsandbytes==0.50.2
python -m pip install 'transformers==5.17.0' 'peft==0.21.0'
hf download TheCulliganMan/kev-27b-nf4 --local-dir kev-27b-nf4
python kev-27b-nf4/patch_kev.py
cd kev-27b-nf4
./serve.sh
The service listens on port 8428 and keeps Kev's kev-latest API model name. serve.sh disables fused weights and CUDA graphs for this NF4 path. Do not merge or cast the adapter into the quantized backbone.
Sources and license
The base is Qwen/Qwen3.8-27B; the decision adapter and head are from jaredpalmer/kev-27b. Both upstream weight releases are Apache-2.0. The evaluation suites and benchmark runner come from the Kev repository. Attribution and source links are provided for the upstream work.
Model tree for TheCulliganMan/kev-27b-nf4
Base model
Qwen/Qwen3.8-27B