Swift-1.5-Qwen3.8-27B-OpenCode-W4A16-AutoRound
A W4A16 quantization (INT4 weights, 16-bit activations) of Swift 1.5,
UkisAI's reasoning-efficient fine-tune of Qwen3.8-27B, calibrated on code (nvidia/OpenCodeInstruct).
It repeats on Swift 1.5 the recipe L00kUp used for their Swift 1.0 OpenCode build (that repository is no longer
available): the same AutoRound code and every one of its settings. VERIFY.txt compares the two runs block by block.
19.0 GB of weights, MTP head included. Tested with vLLM on 2 × RTX 3090 (tensor parallel) and on one A100 80 GB.
In short: on LiveCodeBench v6 it ties ukisai's own W4A16 build problem for problem, and it is 3% faster with a slightly larger KV pool (details below).
What's in the checkpoint (verified from the tensors)
| Part | Precision | Size |
|---|---|---|
| Transformer body: attention, MLP, linear-attention projections | INT4, group 128, symmetric (AutoRound, auto_round:auto_gptq packing) |
12.64 GB |
Linear-attention in_proj_a / in_proj_b, norms, conv, state parameters |
BF16 | 0.06 GB |
| MTP draft head | INT4 layers; mtp.fc and norms BF16 |
0.30 GB |
lm_head |
BF16 | 2.54 GB |
embed_tokens |
BF16 | 2.54 GB |
| Vision tower | BF16 (not quantized) | 0.92 GB |
Quantization recipe
- Source:
ukisai/Swift-1.5-Qwen3.8-27b@bc7a1e10. - Code:
spark-auto-round0.14.3 @2b4136327f4a, the AutoRound fork named in L00kUp's quantization report. The environment is pinned inREQUANT_REQUIREMENTS.lock; every setting is inREQUANT_RECIPE.env. - Calibration:
nvidia/OpenCodeInstruct, 512 samples × 2,048 tokens (packed), batch 4, seed 42. - Tuning: up to 1,000 SignRound iterations per block with best-iteration selection; 46,096 iterations across 64 blocks.
- Scheme: W4A16, group size 128, symmetric. MTP layers quantized. Linear-attention
in_proj_a/in_proj_bandmtp.fckept 16-bit (the fork's defaults for this architecture). - Against L00kUp's Swift 1.0 run (
VERIFY.txt): mean block cosine 0.99639 vs 0.99631, mean PSNR 71.0 vs 70.8 dB, no block notably worse. Per-block figures:quantization-report.csv.
Serving
The quantization method is read from config.json; no --quantization flag is needed. Two configurations were
tested:
vLLM 0.27.1, unpatched (torch 2.13.0+cu130, transformers 5.15.0), one A100 80 GB, MTP off. The LiveCodeBench runs below used this configuration:
vllm serve shuftie/Swift-1.5-Qwen3.8-27B-OpenCode-W4A16-AutoRound \ --dtype bfloat16 --max-model-len 40960 --gpu-memory-utilization 0.92 \ --max-num-batched-tokens 8192 --long-prefill-token-threshold 4096 \ --kv-cache-dtype fp8_e4m3 --trust-remote-code --enable-prefix-caching --enable-chunked-prefill \ --mamba-cache-mode align --prefix-match-unit 16 --reasoning-parser qwen3 \ --default-chat-template-kwargs '{"enable_thinking": true, "reasoning_effort": "xhigh"}'vLLM 0.27.1 with a patch set for Qwen3.8's hybrid attention and MTP (mamba/eagle block, GDN-MTP async speculative ordering, flashinfer decode pin), 2 × RTX 3090, tensor parallel 2, a 262,144-token window,
--speculative-config '{"method":"mtp","num_speculative_tokens":4}'. This is the daily-driver configuration behind the speed figures below. MTP on unpatched vLLM 0.27.1 was not tested.
reasoning_effort accepts xhigh (the default), medium and low, and in this repository also minimal (→ low)
and high / max (→ xhigh). The parent template raised an error on those values.
Evaluation: this build vs ukisai's own W4A16
The comparison arm is ukisai/Swift-1.5-Qwen3.8-27b-W4A16-AutoRound
@ 278de52d (AutoRound 0.15.1, ultrachat_200k calibration, 256 × 2,048 tokens, 200 iterations, BF16 MTP head and
vision). The same weights of Swift 1.5 go into both, so this isolates the quantization recipe.
LiveCodeBench v6, paired
LiveCodeBench's own prompt, code extraction and grader (release_v6), seed 0, thinking xhigh, temperature 1.0 / top_p 0.95 / top_k 20, a 32,768-token budget for thinking and answer together. Configuration 1, both builds on the same 247 hard problems.
| 247 hard problems | this build | ukisai W4A16 |
|---|---|---|
| pass@1 | 46.2% | 45.8% |
| problems solved by only one build | 22 | 21 |
| replies that ran out of budget while still thinking | 107 | 106 |
| pass rate when the thinking finished inside the budget | 81.4% | 80.1% |
| mean completion tokens | 22,742 | 22,483 |
Paired difference +0.4 pp (95% CI −4.9 … +5.7), sign test p = 1.0: the same. The builds also think the same length, problem by problem (paired mean difference +259 tokens, CI −502 … +1,055).
This build alone on 931 of the 1,055 release_v6 problems, weighted to the full set's difficulty mix: 77.2% pass@1 (95% CI 74.9–79.5), with easy 97.8%, medium 88.2% and hard 46.2%. UkisAI's card reports 81.71% for Swift 1.5 in BF16, but with their own harness and five seeds, so the two numbers are not directly comparable. Whatever the gap to BF16 is, the table above shows both 4-bit builds share it; it is not specific to this recipe.
Speed and functional checks (configuration 2, both cards capped at 220 W, back to back)
| this build | ukisai W4A16 | |
|---|---|---|
| decode, tok/s (n = 10 per leg) | 86.5 (4 legs) | 83.9 (2 legs) |
| KV pool, tokens | 556,150 | 541,667 |
| MTP acceptance, decode benchmark | 43–46% (mean length 2.73–2.83) | 45–47% (2.79–2.88) |
| MTP acceptance, thinking and tool traffic | 62.8% (3.51) | 66.9% (3.68) |
| functional gate: 75 scenarios in five packs, incl. multi-step tool calls and reasoning maths | 71/75 | 69/75 |
| GSM-symbolic (30) · AIME hard set (6 problems × 2) | 30/30 · 12/12 | 30/30 · 12/12 |
The decode edge (+3%) is within noise at this sample size. It does not come from better drafts: the stock build's
BF16 MTP head accepts slightly more. The INT4 MTP head here is cheaper per draft step and leaves room for 14.5k more
KV tokens. The two gate failures only the stock build had were a skipped read_file in a tool chain and a reply that
thought until it hit the 16,384-token limit. At n = 75 that difference is not significant (Fisher p ≈ 0.7).
Against L00kUp's Swift 1.0 OpenCode build, on the same checks: gate 71/75 each, GSM-symbolic and AIME perfect for both, decode within noise (86 tok/s).
Known limitations
- Long reasoning needs room. With a 32,768-token budget, 43% of LiveCodeBench hard problems ran out mid-thought, equally for both 4-bit builds. Give it a larger budget if the task allows.
- MTP was tested only on patched vLLM 0.27.1 (configuration 2). Unpatched vLLM was tested with MTP off.
- transformers 5 and the tokenizer: when
tokenizer_config.jsonsaysQwen2Tokenizer, transformers 5 rebuilds the pre-tokenizer from a built-in rule and ignorestokenizer.json's. English, code and accented Latin are unaffected; some scripts, e.g. Hindi and Thai, tokenize differently from training. The tokenizer files here are the parent's, unchanged. The workaround is"tokenizer_class": "TokenizersBackend"intokenizer_config.json. - Sample sizes: the gate is 75 scenarios per build and the LiveCodeBench comparison 247 problems. On hard problems, differences of a few points are not resolved.
- The vision tower is BF16 and unchanged, and was not re-evaluated after quantization.
Credits
UkisAI for Swift 1.5, the Qwen team for Qwen3.8-27B, L00kUp for the OpenCode calibration recipe on Swift 1.0, and Intel's AutoRound.
License
Swift Open License v1.0 (see LICENSE), including its commercial-use threshold (Section 5). Includes the Apache 2.0
license of the Qwen3.8-27B base model (LICENSE-APACHE-2.0) and the upstream NOTICE.
Copyright 2026 UkisAI. Swift Contribution licensed under the Swift Open License v1.0 (https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b/blob/main/LICENSE). Derivative of Qwen3.8-27B, Copyright 2026 Alibaba Cloud, Apache License 2.0.
Changes in this repository (also recorded in NOTICE): the weights of ukisai/Swift-1.5-Qwen3.8-27b were
quantized as described above, config.json was re-saved by the quantizer with quantization_config added, and chat_template.jinja accepts
the reasoning_effort aliases minimal / high / max. Every value the parent template accepted renders
byte-identically.
- Downloads last month
- 64