clef-NVFP4
Unofficial NVFP4 quantization of Cloudflare/clef, the 27B decision model (Qwen3.8-27B backbone) that scores every option of every typed question (choice / score / true-false) in one forward pass. This repo is not affiliated with or endorsed by Cloudflare.
At 17.5 GB (from 55 GB in BF16) it runs on two 16 GB GPUs with tensor parallelism, or on one larger Blackwell
device such as a DGX Spark (GB10). The two big lookup tables (token embeddings and lm_head) are stored as
8-bit with one scale per row, which saves 2.5 GB at no measurable cost in accuracy or speed. It ships with a small
vLLM plugin and a Jev/SystemOne-compatible HTTP server (POST /v1/systemone), because stock vLLM cannot run
Clef's custom joint-schema head.
Updated 2026-10-07: the weight file and the plugin changed together. The token-embedding and
lm_headtables moved out ofmodel.safetensorsintoclef_head/tables_int8.safetensors(8-bit), which made the download 2.5 GB smaller with the same accuracy and speed. If you downloaded this repo before that date:
- Re-download the whole repo and reinstall the plugin (
pip install ./<repo>/vllm_plugin, version 0.2.0 or later). The new weights only work with the new plugin.- Never pair these weights with an older plugin. The checkpoint now carries a version marker, so an older plugin stops at startup with
TypeError: ... unexpected keyword argument 'requires_clef_vllm'. For a few hours on 2026-10-07 the marker was missing: in that state an older plugin starts normally and returns wrong answers, because the two tables stay uninitialized. If you pulled only the weights that day, pull the repo again.- Your old download keeps working as it is, as long as you keep its weights and its plugin together.
Read before using
- Accuracy cost is real: -1.4 points on our 7-task eval (82.8% vs 84.2% BF16; 94.7% top-1 agreement with BF16). That is more than the -0.5 points of our Clef-Flash NVFP4. On these task types, kurcontko/clef-flash-NVFP4 is both more accurate (83.6%) and about 5x faster; pick this 27B model for the workloads where Cloudflare's card shows Clef ahead of Clef-Flash (e.g. GSM8K, CLINC150, RAGTruth), and measure on your own data.
- Text and JSON state only. The vision tower is kept in BF16 but has not been tested; the server rejects
images/videoswith HTTP 400 and runs vLLM with the vision tower disabled.- Use the bundled vLLM plugin, pinned to vLLM 0.28.x (it uses vLLM pooling internals). The checkpoint does not run efficiently in plain
transformers.- Needs the bundled plugin to load at all: the lookup tables are in this repo's own 8-bit format (
clef_head/tables_int8.safetensors), somodel.safetensorsalone is not a complete checkpoint.- Hardware: NVIDIA Blackwell (FP4 tensor cores). Tested on 2x RTX 5070 Ti 16 GB (SM 12.0, PCIe, no NVLink, tensor parallelism 2) and on 1x DGX Spark GB10 (SM 12.1, 128 GB unified memory, single device). Other Blackwell GPUs with 24 GB+ (RTX 5090, RTX PRO 6000, B200) should run it on one device but are untested.
- Evaluated on 1,400 records from 7 public tasks with our own harness, not on Cloudflare's Decision Index.
Quickstart: SystemOne API server
hf download kurcontko/clef-NVFP4 --local-dir clef-NVFP4
docker run --gpus all --ipc=host -p 8000:8000 -v $PWD/clef-NVFP4:/model:ro \
--entrypoint bash vllm/vllm-openai:v0.28.0 -c \
"cp -r /model/vllm_plugin /tmp/p && pip install -q /tmp/p && clef-systemone --model /model --tp 2"
On one large GPU (DGX Spark, RTX 5090, B200) drop --tp 2. With --tp 2 the server turns FlashInfer autotuning off (vLLM 0.28 can deadlock autotuning NVFP4 kernels across
tensor-parallel ranks); --flashinfer-autotune on|off overrides that. It warms up its kernels before /health reports
ready (a few minutes from a cold container).
curl localhost:8000/v1/systemone -H 'content-type: application/json' -d '{
"model": "clef",
"state": "Our checkout started returning errors and orders are blocked.",
"questions": {
"department": {"type": "choice", "instructions": "Which team should handle the message?",
"criteria": {"billing": "Payments or invoices", "technical": "Bugs or outages"}},
"urgency": {"type": "score", "criteria": ["Can wait", "This week", "Today"]},
"outage": {"type": "noul", "instructions": "Is a service down?"}}}'
Request and response bodies match the systemone() function of the original release (answers keyed by question ID:
choice with confidence and probabilities, score with the expected score and legend, noul with the
probability of true; plus usage). Also served: GET /v1/models, GET /health.
Quickstart: Python (offline vLLM)
# pip install ./clef-NVFP4/vllm_plugin (in an environment with vllm==0.28.*)
import sys
from transformers import AutoTokenizer
from vllm import LLM, PoolingParams
from clef_vllm import ARCHITECTURE, TASK, schema_layout
path = "./clef-NVFP4"
sys.path.insert(0, path)
from joint_schema_model import encode_record
def main():
# Keep `llm` local: vLLM shuts its engine process down when the LLM object is freed.
llm = LLM(model=path, hf_overrides={"architectures": [ARCHITECTURE]}, runner="pooling",
tensor_parallel_size=2, kernel_config={"enable_flashinfer_autotune": False},
enable_prefix_caching=False, language_model_only=True, max_model_len=16384)
tokenizer = AutoTokenizer.from_pretrained(path)
record = {
"state": {"invoice": {"vendor": "Acme", "total": 1250.0, "currency": "USD", "status": "overdue"}},
"questions": {
"status": {"type": "choice", "instructions": "What is the invoice status?",
"criteria": {"paid": "Invoice is paid.", "overdue": "Invoice is past due.", "draft": "Not sent."}},
"large": {"type": "noul", "instructions": "Is the total above 1000 USD?"},
},
}
encoded = encode_record(tokenizer, record)
params = PoolingParams(task=TASK, extra_kwargs={"clef_questions": schema_layout(encoded)})
output = llm.encode([{"prompt_token_ids": list(encoded.input_ids)}], pooling_params=[params], pooling_task=TASK)[0]
logits, offset = output.outputs.data.float(), 0 # every option logit of every question, in order
for question in encoded.questions:
n = len(question.option_ids)
print(question.question_id, dict(zip(question.option_ids, logits[offset:offset + n].softmax(-1).tolist())))
offset += n
if __name__ == "__main__":
main()
Results
Accuracy and agreement with BF16
1,400 records, 200 per task, from held-out splits: BANKING77 (test), ARC-Challenge (test), MMLU (test), HellaSwag
(validation), ANLI R3 (test), BoolQ (validation), and a multi-question ANLI R1 set (3 questions per record). Acc is
against gold labels; agree is top-1 agreement with BF16; KL is mean KL(BF16 ‖ NVFP4) per question. BF16 is the
original release in transformers 5.10.2 (run on a DGX Spark); NVFP4 was measured in vLLM with real FP4 kernels.
| Task | BF16 acc | NVFP4 acc | Agree | KL |
|---|---|---|---|---|
| BANKING77 (77-way intent) | 96.0 | 95.5 | 99.5 | 0.011 |
| ARC-Challenge | 98.5 | 97.0 | 98.5 | 0.012 |
| MMLU | 93.0 | 92.0 | 95.5 | 0.022 |
| HellaSwag | 97.5 | 97.0 | 98.5 | 0.006 |
| ANLI R3 | 51.0 | 47.5 | 88.5 | 0.046 |
| BoolQ | 90.5 | 90.5 | 99.0 | 0.009 |
| Multi-question ANLI R1 | 73.5 | 71.2 | 90.8 | 0.030 |
| All | 84.2 | 82.8 | 94.7 | 0.022 |
The table was measured before the lookup tables were made 8-bit; with them, three more runs gave 94.3-94.7% agreement, KL 0.019-0.021 and 82.6-82.8% accuracy, the same within run-to-run noise.
A simulated NVFP4 run in transformers (weights dequantized, activations fake-quantized) gave 82.6% / 94.6% / 0.021,
matching vLLM, so the gap is quantization error, not the serving path. For comparison on the same set, the original
Clef-Flash scores 84.1% in BF16 and 83.6% as NVFP4.
Throughput
Same 1,400 records (717k tokens, ~512 per record).
| Setup | Hardware | Records/s | Tokens/s |
|---|---|---|---|
BF16, transformers 5.10.2 (original release code) |
1x DGX Spark (GB10) | 2.0 | 1,002 |
| NVFP4, vLLM + plugin, single device | 1x DGX Spark (GB10) | 4.2 | 2,139 |
| NVFP4, vLLM + plugin, tensor parallel 2 | 2x RTX 5070 Ti 16 GB | 8.4 | 4,288 |
On the GB10 the NVFP4 model agreed with BF16 on 94.1% of questions (KL 0.022, accuracy 82.7%), the same picture as on
the RTX 5070 Ti pair; it also leaves room for ~747k tokens of KV cache at --gpu-memory-utilization 0.6. The GB10
rows were measured before the lookup tables were made 8-bit; on the RTX 5070 Ti pair that change left throughput
unchanged (4,226-4,303 tok/s).
Through the HTTP server on the 2x RTX 5070 Ti pair (clean install from this repo, default settings): 64 concurrent clients 8.2 req/s, 4,209 tok/s (p50 4.5 s, p95 27 s, queueing-bound); 8 concurrent clients 8.1 req/s, p50 565 ms, p95 3.4 s. Answers agreed with BF16 on 94.6 to 95.4% of questions.
Single-request latency
One request at a time over HTTP, end to end, on the 2x RTX 5070 Ti pair with --tp 2 --gdn-backend flashinfer
(median of 80):
| Request | Input tokens | Median | p95 |
|---|---|---|---|
| One true/false question | 147 | 62 ms | 63 ms |
| One 3-option choice | 242 | 78 ms | 79 ms |
| 3 questions (the example above) | 300 | 94 ms | 95 ms |
The bundled server captures CUDA graphs for prompt-sized steps (vLLM's default stops at 2 x max_num_seqs tokens and
leaves prefill-only requests on a slow eager path) and can force FlashInfer's linear-attention kernel on RTX 50-series
with --gdn-backend flashinfer (experimental; see the
clef-flash-NVFP4 card for both). Responses carry a Server-Timing
header.
Quantization details
- Tool: llm-compressor 0.14.0 (compressed-tensors 0.19.0),
QuantizationModifier(targets="Linear", scheme="NVFP4"), run on a DGX Spark in about 23 minutes. - Scheme: NVFP4 W4A4: FP4 (E2M1) weights and activations, 16-element groups with FP8 (E4M3) group scales and one FP32 global scale per tensor; activation global scales calibrated, group scales dynamic.
- Calibration: 672 records (96 from the train split of each eval task, encoded exactly as at inference).
- Lookup tables in 8-bit: Clef never multiplies by its token-embedding or
lm_headmatrices: embeddings are looked up by token id, and the joint head only reads thelm_headrows of each request's option tokens. Both are stored as int8 with one scale per row (scripts/shrink_tables.py): relative error 0.9% and 1.2%, mean row cosine 0.99996 and 0.99992. - Kept in BF16: the vision tower, the Gated DeltaNet gate projections
linear_attn.in_proj_a/in_proj_b, norms, and the joint schema head. - Gated DeltaNet fused scale: vLLM fuses
linear_attn.in_proj_qkvandin_proj_zinto one GEMM, which needs a shared NVFP4 global scale;scripts/quantize.pyadds that pair to llm-compressor's fused groups (all 48 GDN layers verified). Without it, vLLM silently mis-scales the qkv weights. recipe.yamlis the exact recipe llm-compressor saved. Post-training quantization only: no quantization-aware training.
How the vLLM plugin works
vllm_plugin/ (package clef-vllm) registers ClefFlashForDecision, a pooling-model subclass of vLLM's
Qwen3_5ForConditionalGeneration (the same class serves Clef and Clef-Flash). Its pooler collects each request's final
hidden states across chunked prefill and runs the original JointSchemaHead (from joint_schema_model.py) on the full
prompt; each request carries its question/option token spans in PoolingParams.extra_kwargs["clef_questions"].
The 8-bit embedding table sits on every GPU; the 8-bit lm_head table is only read by the pooler, so under tensor
parallelism it stays in host memory on rank 0 and only each request's option-token rows are copied to the GPU. That
leaves about 2.8 GiB of KV cache per 16 GB card at the default settings.
Files
| Path | Purpose |
|---|---|
model.safetensors, config.json, recipe.yaml |
Quantized backbone without the two lookup tables (14.7 GB, compressed-tensors format) |
clef_head/tables_int8.safetensors |
Token embeddings and lm_head as int8 + per-row scale (2.5 GB) |
clef_head/joint_head* |
Joint schema head, BF16; weights unchanged from the original release, config has one added key (requires_clef_vllm). In a subfolder so vLLM does not load it as backbone weights |
joint_schema_model.py |
Original release code: encode_record, systemone_answer, the head module (unchanged) |
tokenizer*.json, chat_template.jinja, processor_config.json, generation_config.json |
Tokenizer and processor (unchanged) |
vllm_plugin/ |
vLLM plugin and clef-systemone server |
scripts/ |
build_data.py (calibration/eval sets), quantize.py and shrink_tables.py (how this checkpoint was made) |
License and attribution
Apache-2.0, following Cloudflare/clef (post-trained from Qwen/Qwen3.8-27B). All credit for the model goes to its original authors; this repository only adds the quantization, the vLLM plugin and the server. Not an official Cloudflare release.
- Downloads last month
- 23