Swift-1.5-Qwen3.8-27B-OpenCode-W4A16-AutoRound

A W4A16 quantization (INT4 weights, 16-bit activations) of Swift 1.5, UkisAI's reasoning-efficient fine-tune of Qwen3.8-27B, calibrated on code (nvidia/OpenCodeInstruct).

It repeats on Swift 1.5 the recipe L00kUp used for their Swift 1.0 OpenCode build (that repository is no longer available): the same AutoRound code and every one of its settings. VERIFY.txt compares the two runs block by block.

19.0 GB of weights, MTP head included. Tested with vLLM on 2 × RTX 3090 (tensor parallel) and on one A100 80 GB.

In short: on LiveCodeBench v6 it ties ukisai's own W4A16 build problem for problem, and it is 3% faster with a slightly larger KV pool (details below).

What's in the checkpoint (verified from the tensors)

Part Precision Size
Transformer body: attention, MLP, linear-attention projections INT4, group 128, symmetric (AutoRound, auto_round:auto_gptq packing) 12.64 GB
Linear-attention in_proj_a / in_proj_b, norms, conv, state parameters BF16 0.06 GB
MTP draft head INT4 layers; mtp.fc and norms BF16 0.30 GB
lm_head BF16 2.54 GB
embed_tokens BF16 2.54 GB
Vision tower BF16 (not quantized) 0.92 GB

Quantization recipe

  • Source: ukisai/Swift-1.5-Qwen3.8-27b @ bc7a1e10.
  • Code: spark-auto-round 0.14.3 @ 2b4136327f4a, the AutoRound fork named in L00kUp's quantization report. The environment is pinned in REQUANT_REQUIREMENTS.lock; every setting is in REQUANT_RECIPE.env.
  • Calibration: nvidia/OpenCodeInstruct, 512 samples × 2,048 tokens (packed), batch 4, seed 42.
  • Tuning: up to 1,000 SignRound iterations per block with best-iteration selection; 46,096 iterations across 64 blocks.
  • Scheme: W4A16, group size 128, symmetric. MTP layers quantized. Linear-attention in_proj_a/in_proj_b and mtp.fc kept 16-bit (the fork's defaults for this architecture).
  • Against L00kUp's Swift 1.0 run (VERIFY.txt): mean block cosine 0.99639 vs 0.99631, mean PSNR 71.0 vs 70.8 dB, no block notably worse. Per-block figures: quantization-report.csv.

Serving

The quantization method is read from config.json; no --quantization flag is needed. Two configurations were tested:

  1. vLLM 0.27.1, unpatched (torch 2.13.0+cu130, transformers 5.15.0), one A100 80 GB, MTP off. The LiveCodeBench runs below used this configuration:

    vllm serve shuftie/Swift-1.5-Qwen3.8-27B-OpenCode-W4A16-AutoRound \
      --dtype bfloat16 --max-model-len 40960 --gpu-memory-utilization 0.92 \
      --max-num-batched-tokens 8192 --long-prefill-token-threshold 4096 \
      --kv-cache-dtype fp8_e4m3 --trust-remote-code --enable-prefix-caching --enable-chunked-prefill \
      --mamba-cache-mode align --prefix-match-unit 16 --reasoning-parser qwen3 \
      --default-chat-template-kwargs '{"enable_thinking": true, "reasoning_effort": "xhigh"}'
    
  2. vLLM 0.27.1 with a patch set for Qwen3.8's hybrid attention and MTP (mamba/eagle block, GDN-MTP async speculative ordering, flashinfer decode pin), 2 × RTX 3090, tensor parallel 2, a 262,144-token window, --speculative-config '{"method":"mtp","num_speculative_tokens":4}'. This is the daily-driver configuration behind the speed figures below. MTP on unpatched vLLM 0.27.1 was not tested.

reasoning_effort accepts xhigh (the default), medium and low, and in this repository also minimal (→ low) and high / max (→ xhigh). The parent template raised an error on those values.

Evaluation: this build vs ukisai's own W4A16

The comparison arm is ukisai/Swift-1.5-Qwen3.8-27b-W4A16-AutoRound @ 278de52d (AutoRound 0.15.1, ultrachat_200k calibration, 256 × 2,048 tokens, 200 iterations, BF16 MTP head and vision). The same weights of Swift 1.5 go into both, so this isolates the quantization recipe.

LiveCodeBench v6, paired

LiveCodeBench's own prompt, code extraction and grader (release_v6), seed 0, thinking xhigh, temperature 1.0 / top_p 0.95 / top_k 20, a 32,768-token budget for thinking and answer together. Configuration 1, both builds on the same 247 hard problems.

247 hard problems this build ukisai W4A16
pass@1 46.2% 45.8%
problems solved by only one build 22 21
replies that ran out of budget while still thinking 107 106
pass rate when the thinking finished inside the budget 81.4% 80.1%
mean completion tokens 22,742 22,483

Paired difference +0.4 pp (95% CI −4.9 … +5.7), sign test p = 1.0: the same. The builds also think the same length, problem by problem (paired mean difference +259 tokens, CI −502 … +1,055).

This build alone on 931 of the 1,055 release_v6 problems, weighted to the full set's difficulty mix: 77.2% pass@1 (95% CI 74.9–79.5), with easy 97.8%, medium 88.2% and hard 46.2%. UkisAI's card reports 81.71% for Swift 1.5 in BF16, but with their own harness and five seeds, so the two numbers are not directly comparable. Whatever the gap to BF16 is, the table above shows both 4-bit builds share it; it is not specific to this recipe.

Speed and functional checks (configuration 2, both cards capped at 220 W, back to back)

this build ukisai W4A16
decode, tok/s (n = 10 per leg) 86.5 (4 legs) 83.9 (2 legs)
KV pool, tokens 556,150 541,667
MTP acceptance, decode benchmark 43–46% (mean length 2.73–2.83) 45–47% (2.79–2.88)
MTP acceptance, thinking and tool traffic 62.8% (3.51) 66.9% (3.68)
functional gate: 75 scenarios in five packs, incl. multi-step tool calls and reasoning maths 71/75 69/75
GSM-symbolic (30) · AIME hard set (6 problems × 2) 30/30 · 12/12 30/30 · 12/12

The decode edge (+3%) is within noise at this sample size. It does not come from better drafts: the stock build's BF16 MTP head accepts slightly more. The INT4 MTP head here is cheaper per draft step and leaves room for 14.5k more KV tokens. The two gate failures only the stock build had were a skipped read_file in a tool chain and a reply that thought until it hit the 16,384-token limit. At n = 75 that difference is not significant (Fisher p ≈ 0.7).

Against L00kUp's Swift 1.0 OpenCode build, on the same checks: gate 71/75 each, GSM-symbolic and AIME perfect for both, decode within noise (86 tok/s).

Known limitations

  • Long reasoning needs room. With a 32,768-token budget, 43% of LiveCodeBench hard problems ran out mid-thought, equally for both 4-bit builds. Give it a larger budget if the task allows.
  • MTP was tested only on patched vLLM 0.27.1 (configuration 2). Unpatched vLLM was tested with MTP off.
  • transformers 5 and the tokenizer: when tokenizer_config.json says Qwen2Tokenizer, transformers 5 rebuilds the pre-tokenizer from a built-in rule and ignores tokenizer.json's. English, code and accented Latin are unaffected; some scripts, e.g. Hindi and Thai, tokenize differently from training. The tokenizer files here are the parent's, unchanged. The workaround is "tokenizer_class": "TokenizersBackend" in tokenizer_config.json.
  • Sample sizes: the gate is 75 scenarios per build and the LiveCodeBench comparison 247 problems. On hard problems, differences of a few points are not resolved.
  • The vision tower is BF16 and unchanged, and was not re-evaluated after quantization.

Credits

UkisAI for Swift 1.5, the Qwen team for Qwen3.8-27B, L00kUp for the OpenCode calibration recipe on Swift 1.0, and Intel's AutoRound.

License

Swift Open License v1.0 (see LICENSE), including its commercial-use threshold (Section 5). Includes the Apache 2.0 license of the Qwen3.8-27B base model (LICENSE-APACHE-2.0) and the upstream NOTICE.

Copyright 2026 UkisAI. Swift Contribution licensed under the Swift Open License v1.0 (https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b/blob/main/LICENSE). Derivative of Qwen3.8-27B, Copyright 2026 Alibaba Cloud, Apache License 2.0.

Changes in this repository (also recorded in NOTICE): the weights of ukisai/Swift-1.5-Qwen3.8-27b were quantized as described above, config.json was re-saved by the quantizer with quantization_config added, and chat_template.jinja accepts the reasoning_effort aliases minimal / high / max. Every value the parent template accepted renders byte-identically.

Downloads last month
64
Safetensors
Model size
3B params
Tensor type
BF16
·
I32
·
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for shuftie/Swift-1.5-Qwen3.8-27B-OpenCode-W4A16-AutoRound

Base model

Qwen/Qwen3.8-27B
Quantized
(52)
this model