Swift-1.5-Qwen3.8-27b-FP8

ukisai/Swift-1.5-Qwen3.8-27b quantized to FP8 with the same recipe and checkpoint format as the official Qwen/Qwen3.8-27B-FP8: weights in FP8 E4M3 with 128x128 block scales, activations dynamically quantized (activation_scheme: dynamic), serialized in the compressed-tensors/ quant_method: fp8 layout that vLLM loads directly.

Quantization

property value
weight format FP8 E4M3, per-128x128-block scales (weight_scale_inv, BF16)
activation format FP8 E4M3, dynamic (no stored input scales)
quantized linears 407 (400 language-decoder + 7 MTP)
unquantized vision tower (model.visual.*), lm_head, norms, embeddings, GDN conv/state parameters

modules_to_not_convert is a superset of the reference checkpoint's list (1111 vs 882 entries): every module that is BF16 in this checkpoint is listed, and the quantized tensor set is exactly the reference's.

Verification

  • Tensor inventory matched against Qwen/Qwen3.8-27B-FP8: identical 1606 keys (407 F8_E4M3 + 1199 BF16), identical quantized module set, no name or shape differences.
  • Served with vLLM Radiance (vLLM 0.28.0, gfx1201, tensor-parallel 2, R4D attention, prefix caching) and produced correct generations.

Evaluation

Measured with the gsm8k bench of check-source, the successor to gsm8k-eval. It reproduces the lm-evaluation-harness gsm8k task exactly: identical prompt format, per-document 5-shot sampling (seed 1234), filters and generation settings for every checkpoint, full 1319-item test split. check-source also carries the chat-template GSM8K protocol (one user turn, reasoning_effort levels) plus AIME and MATH-500 benches.

Raw result files for these runs: ethantodd4l/check-source-results.

checkpoint thinking strict thinking flexible non-thinking strict non-thinking flexible
this model (W8A8 FP8, fp8 KV) 85.75% 85.67% ±1.89 85.97% 88.70% ±1.71
Swift-1.5-Qwen3.8-27b bf16 (base) pending pending pending pending

Chat protocol (check-source gsm8k-chat)

The raw bench fixes the encoder prompt and the two AMD modes; the chat bench puts the same lm-eval 5-shot text in a single user turn and lets the checkpoint's template choose the thinking level, with the reasoning-level sampling defaults (T=1.0, top_p=0.95, top_k=20). Same checkpoint family, only the serving configuration changes (200 items, seed 1234, target-only, no speculation):

weights activations KV cache none medium xhigh mean tokens none/med/xhigh
bf16 (base) bf16 fp16 97.5% 98.5% 98.5% 178 / 328 / 427
FP8 W8A8 (this model) fp8 dynamic per-token fp16 98.5% 98.0% 99.5% 162 / 328 / 452
FP8 W8A8 (this model) fp8 dynamic per-token fp8 97.0% 98.0% 98.0% 168 / 305 / 481
MXFP4 RTN fp8 WMMA (W4A8) fp16 96.5% 98.0% 97.0% 148 / 317 / 500
MXFP4 RTN fp8 WMMA (W4A8) fp8 96.5% 99.0% 96.5% 162 / 308 / 497

Scores are flexible-extract; at 200 items the 95% interval is roughly 1.5 points, so the accuracy differences above sit inside noise. The consistent signals are the KV direction (fp16 beats fp8 in every pair) and the token counts. A full 1319-item run plus AIME/MATH-500 is the next step.

Activation and KV-cache precision

Dequantized to bf16 (undoing the 128x128 block scales), this checkpoint is nearly lossless: next-token logits for "The capital of France is" against the bf16 base give rel=0.024, corr=0.999 and the same top-1 token (Paris). The GSM8K gap above therefore comes from the runtime numerics, not the stored weights:

  • FP8 activations (W8A8 dynamic, AITER per-token) are the dominant accuracy cost on this stack. Weight-only (W8A16) would remove it, but this vLLM build has no ROCm W8A16 kernel; it needs a radiance-side kernel. The ~9-point gap above is the full-split raw protocol; on the first 200 items with the chat protocol this same W8A8 configuration is within noise of the bf16 base, so the size of the gap is protocol- and level-dependent.
  • fp8 KV cache costs about 1.3 points by itself in a 300-item seeded A/B (80.0% -> 81.3% flexible). Switching to fp16 KV (--kv-cache-dtype auto) does not recover the bulk of the gap.
  • Production guidance: use fp16 KV for accuracy (at ~half the KV capacity), and prefer the MXFP4 RTN sibling for this base model until a W8A16 kernel exists.

Serving

vllm serve ethantodd4l/Swift-1.5-Qwen3.8-27b-FP8 \
  --served-model-name swift-1.5-fp8 \
  --tensor-parallel-size 2 \
  --max-model-len 16384 \
  --kv-cache-dtype auto

The FP8 lane is the default of vLLM Radiance (ROCm/gfx1201, vLLM 0.28.0); the MXFP4-specific switches (RADIANCE_MXFP4, RADIANCE_MXFP4_W4A8, RADIANCE_QUARK_BF16_MTP) do not apply to this checkpoint. The evaluation row above was measured with fp8 KV; auto (fp16) recovers about 1.3 points in a 300-item probe, though not the bulk of the W8A8 gap.

License

The quantization is a derivative of ukisai/Swift-1.5-Qwen3.8-27b and is distributed under the Swift Open License v1.0 (see LICENSE). The Qwen3.8-27B base model and the tokenizer/config files retained from it remain under their original Apache License 2.0 terms; see NOTICE.

Downloads last month
53
Safetensors
Model size
28B params
Tensor type
BF16
·
F8_E4M3
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ethantodd4l/Swift-1.5-Qwen3.8-27b-FP8

Base model

Qwen/Qwen3.8-27B
Quantized
(52)
this model