Swift-1.5-Qwen3.8-27B-NVFP4 GGUF (Q8mix)

A llama.cpp GGUF conversion of ukisai/Swift-1.5-Qwen3.8-27b-NVFP4, UkisAI's calibrated ModelOpt NVFP4 + FP8 checkpoint of Swift-1.5-Qwen3.8-27B. The 193 NVFP4 tensors are not requantized — they are repacked into GGML's native NVFP4 super-block layout with their original per-16 UE4M3 scales and per-tensor weight_scale_2 factors preserved bit-for-bit.

File Size BPW
Swift-1.5-Qwen3.8-27B-NVFP4-Q8mix.gguf 19.74 GB 5.78 (27.32 B params)
mmproj-Swift-1.5-Qwen3.8-27B-F16.gguf 885 MB official UkisAI projector
Swift-1.5-Qwen3.8-27B-NVFP4-Q8mix.tensor-types.txt 41 KB exact per-tensor layout used by llama-quantize

Tensor layout

Tensors Type Source
MLP ffn_gate/ffn_up/ffn_down for all 64 layers + output (lm_head) — 193 tensors NVFP4 ModelOpt NVFP4, byte-identical to the checkpoint (incl. 386 scale tensors)
Linear-attention attn_qkv, attn_gate, ssm_out (48 layers); full-attention attn_q/k/v/output (16 layers) — 208 tensors Q8_0 FP8 (E4M3) in the checkpoint, dequantized then Q8_0
token_embd Q8_0 BF16 in the checkpoint
MTP head (blk.64, 15 tensors) Q6_K attn / Q4_K FFN, F32 norms unsloth/Qwen3.8-27B-GGUF MTP/mtp-Qwen3.8-27B-Q4_0.gguf (UkisAI did not fine-tune the MTP head)
ssm_alpha, ssm_beta, ssm_conv1d, norms, per-tensor scales F32 BF16 in the checkpoint (lossless)

The MTP head is included (qwen35.nextn_predict_layers = 1), so --spec-type draft-mtp works as a self-speculative draft.

How it was made

llama.cpp commit 2145525a4081d66ff1a87cf43ef809f95a85ac0c (2026-09-26):

  1. convert_hf_to_gguf.py --outtype bf16 --fp8-as-q8 — the converter natively repacks the ModelOpt NVFP4 group into GGML NVFP4 and turns the FP8 group into Q8_0.
  2. llama-quantize --tensor-type-file <the included .tensor-types.txt> with base type q8_0. NVFP4 and Q8_0 tensors map to their own type and are copied byte-for-byte; the BF16 leftovers become Q8_0; recurrent gates stay F32.
  3. blk.64.* (MTP head) spliced byte-for-byte from unsloth's Q4_0 MTP file with a byte-level GGUF rewrite (swap_mtp.py in the build workspace): metadata and all other 1237 tensors untouched, 185 MB smaller.

Verification

  • 1147/1252 tensors byte-identical between the raw conversion and the final file (everything llama-quantize copied).
  • All 193 NVFP4 tensors independently repacked from the source safetensors and compared byte-for-byte: exact match.
  • GGML nvfp4 decode vs ModelOpt weight * weight_scale * weight_scale_2: exact (max abs diff 0.0 on sampled tensors).
  • Metadata/tensor names vs UkisAI's own GGUF: all 41 metadata keys match (except cosmetic general.version), names = official set + 386 scale companions, all shapes identical.
  • Inference smoke test on RTX PRO 4000 Blackwell (24 GB): 182.6 t/s prompt, 20.0 t/s generation; coherent output.
  • MTP splice: metadata identical, 15/15 blk.64.* tensors byte-identical to unsloth's file, all other 1237 tensors byte-identical to the first build.

MTP / speculative decoding benchmark

RTX PRO 4000 Blackwell (24 GB), 3 prompts x 256 tokens, greedy (temperature 0), --spec-type draft-mtp:

MTP head gen t/s speedup acceptance mean accepted
no spec 19.8 1.00x - -
NVFP4 (RTN, all 8 weights) 40.2 2.03x 72.0% 3.16/4
Q8_0 (original build) 41.0 2.07x 73.7% 3.21/4
Q6_K attn / Q4_K FFN (this file) 41.3 2.09x 73.7% 3.21/4

--spec-draft-n-max 3 is the optimum; asking for more drafts costs throughput (4/5/6 drafts: 40.5 / 38.1 / 37.2 t/s) because acceptance falls with position faster than the extra drafts pay off.

An RTN-NVFP4 MTP head (all 8 weights, no calibration) was also built and measured interleaved at n_max 3/4/5, two runs each: it is equal or behind the Q4_K/Q6_K head at every n_max (40.3 / 39.8 / 37.1 t/s average; acceptance 72.0 / 64.4 / 55.6%), so the K-quant head is kept.

Usage

llama-server -m Swift-1.5-Qwen3.8-27B-NVFP4-Q8mix.gguf \
  --mmproj mmproj-Swift-1.5-Qwen3.8-27B-F16.gguf \
  --spec-type draft-mtp --spec-draft-n-max 3 \
  -c 262144 -fa on --jinja

NVFP4 runs natively on NVIDIA Blackwell. Chat template: reasoning_effort accepts xhigh (default), medium, low. Sampling: temp 1.0, top_p 0.95, top_k 20 (stored in the GGUF).

License

Swift weights are under the Swift Open License v1.0 (see the source model card). Qwen3.8-27B is Apache 2.0. All credit for the model and the NVFP4 calibration goes to UkisAI.

Downloads last month
854
GGUF
Model size
27B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for BroLaurens/Swift-1.5-Qwen3.8-27B-NVFP4-GGUF

Base model

Qwen/Qwen3.8-27B
Quantized
(3)
this model