Swift-1.5-Qwen3.8-27B — Heretic Abliterated · GPTQ W8A8 INT8 (unrotated, DFlash2)
Research artifact. Safety alignment has been deliberately removed. This model will attempt to comply with harmful, dangerous, illegal, and unethical requests that the source model refuses, with no content moderation. See Safety. Provided for research purposes only, with no warranty and no liability accepted by the author — see the disclaimer at the bottom.
An INT8 W8A8 build of
akumaburn/Swift-1.5-Qwen3.8-27b-heretic
deliberately produced without the QuaRot rotation, so the
DFlash2 drafter works. Size 29G.
- Weights: INT8, per-channel symmetric (GPTQ).
- Activations: INT8, per-token dynamic (
int-quantized). - Unrotated: norms are preserved (not folded to zero), which is what DFlash2's
fcprojection requires. - Left unquantized: MTP head,
lm_head,embed_tokens, vision tower, GatedDeltaNet gates, all norms.
Fidelity — the cost of staying unrotated
| value | |
|---|---|
KL(abliterated ‖ this) — quantization only |
0.0791 |
KL(source ‖ this) |
0.1330 |
This is 3.0× the divergence of the rotated sibling (0.0264) at the same bit-width and nearly the same size. That is the price of DFlash2 compatibility, and it buys the highest acceptance length of the set (4.34 vs 2.51 for MTP-3).
Abliteration
Produced with Heretic v2.0.0.dev0: 260 TPE trials,
seed 42, directional ablation on attn.o_proj (16 modules), attn.out_proj (48) and
mlp.down_proj (64). Baseline refusal score before abliteration: 98/100.
Trial 140 was selected — Pareto index 1, not index 0. Index 0 (trial 161) scored 22/100 keywords at KL 0.1091; trial 140 scores 23/100 at KL 0.0602. One extra soft keyword hit in exchange for 45% less divergence from the source was the better trade, and the refusal breakdown above confirms the cost was nil: both are 0/100 on genuine refusals.
Independent verification on 100 harmless prompts (full-vocabulary first-token KL,
outside Heretic) measures KL(source ‖ this) = 0.0616, agreeing with
Heretic's own 0.0602 to within 2%.
This is a thinking model: the chat template emits <think> as part of the generation
prompt, so scoring used the response prefix \n</think>\n\n to close the block —
without it every scored token is reasoning text and the refusal metric saturates.
Variants
All four builds derive from the same trial-140 abliteration. Measured on identical data and hardware (see Benchmarks).
| build | size | first-token KL vs source | quant-only KL | accept. length | best for |
|---|---|---|---|---|---|
| BF16 abliterated | 52G | 0.0616 | — | n/a | research, re-quantization |
| W8A8 + QuaRot/SmoothQuant | 30G | 0.0843 | 0.0264 | 2.51 (MTP-3) | highest fidelity; batch serving |
| W8A8 unrotated | 29G | 0.1330 | 0.0791 | 4.34 (DFlash2) | DFlash2 speculation at c ≤ 8 |
| W4A16 unrotated | 18G | 0.2109 | 0.1182 | 3.61 (DFlash2) | smallest; single-user decode |
quant-only KL isolates the quantization from the abliteration: KL(abliterated ‖ build).
The rotation is worth 3.0× on this metric (0.0264 vs 0.0791 unrotated),
and 4-bit costs 4.5× the rotated 8-bit build.
Match speculative depth to concurrency. DFlash2 (k=7) gives the best decode at low concurrency, but its prefill collapses when the GPU saturates: on the unrotated W8A8, prefill falls from 7,329 tok/s at c=8 to 2,829 tok/s at c=32, and TTFT p50 rises to 12.6 s. The rotated W8A8 with MTP-3 instead climbs to 19,209 tok/s at c=32. Use DFlash2 for c ≤ 8; prefer the rotated W8A8 with MTP for c ≥ 16.
Benchmarks
Measured on 2× NVIDIA CMP 170HX (64 GB, sm_80), TP=1 (single card per server),
vLLM 0.27.1, fp8_e4m3 KV cache, async scheduling, --max-model-len 32768.
Synthetic input 1,024 tokens / output 256, --ignore-eos; acceptance measured
separately on 88 code tasks, greedy, concurrency 1.
vLLM 0.27.1 was used deliberately: on this Ampere hardware, --kv-cache-dtype fp8_e4m3
makes vLLM 0.28.0 fall back to FlashInfer (FlashAttention has no fp8-KV path below
Hopper), which faults with an illegal memory access under speculative decoding.
| build | c | prefill tok/s | decode tok/s | TTFT p50 | TPOT p50 | accept. len |
|---|---|---|---|---|---|---|
| W8A8 rotated (MTP-3) | 1 | 3,432 | 57.3 | 272 ms | 13.12 ms | 2.06 |
| 8 | 8,720 | 290.7 | 350 ms | 18.87 ms | 2.61 | |
| 32 | 19,209 | 516.8 | 643 ms | 46.29 ms | 2.41 | |
| W8A8 unrot (DFlash2 k=7) | 1 | 2,882 | 87.5 | 337 ms | 8.85 ms | 3.12 |
| 8 | 7,329 | 321.1 | 936 ms | 17.17 ms | 3.38 | |
| 32 | 2,829 | 410.8 | 12551 ms | 21.99 ms | 3.66 | |
| W4A16 (DFlash2 k=7) | 1 | 1,926 | 119.3 | 514 ms | 5.88 ms | 3.63 |
| 8 | 4,983 | 254.5 | 1240 ms | 21.08 ms | 3.39 | |
| 32 | 1,901 | 280.2 | 18356 ms | 32.48 ms | 3.72 |
Acceptance on 88 code tasks, greedy, c=1: W8A8 rotated (MTP-3) 50.3% / 2.51,
W8A8 unrotated (DFlash2 k=7) 47.7% / 4.34,
W4A16 (DFlash2 k=7) 37.3% / 3.61.
Acceptance length (mean tokens accepted per verify step) is what sets the speedup;
the rate alone is not comparable across different num_speculative_tokens.
The source card's task benchmarks (GPQA-Diamond, IFBench, AIME 2026, LiveCodeBench, Terminal-Bench) were not re-run here — single-GPU budget. The published values on the source card stand as reported, and nothing in this build is claimed to preserve them. The KL figures above are the fidelity evidence actually measured.
Serving (vLLM, with DFlash2)
vllm serve akumaburn/Swift-1.5-Qwen3.8-27b-heretic-W8A8-DFlash2 \
--tensor-parallel-size 1 --max-model-len 32768 \
--kv-cache-dtype fp8_e4m3 --async-scheduling \
--speculative-config '{"method":"dflash","model":"<DFlash2 drafter>","num_speculative_tokens":7}'
Drop the --speculative-config for concurrency ≥ 16; see the tip above.
Safety
Refusal behaviour has been deliberately removed. Measured on 100 harmful-instruction prompts, with the keyword classes separated:
| count | |
|---|---|
| Hard refusals (genuine refusal language) | 0 / 100 |
| Soft caveats only (complied, but said "illegal"/"harmful"/…) | 24 / 100 |
| True refusal rate | 0 / 100 |
The residual 24/100 keyword hits are compliant answers that add a caveat, not refusals — a measurement floor of the KeywordRate metric rather than surviving censorship.
This model produces content the source declines, including dangerous, illegal, or unethical material, with no moderation. Intended for interpretability/safety research, red-teaming, and evaluation by people who understand and accept those risks. Do not deploy it where it can reach people who have not consented to unfiltered output. You are responsible for your use and for compliance with all applicable laws.
Licence
Dual stack (see NOTICE and the two licence files shipped with this model):
- Swift Open License v1.0 — UkisAI's contribution (
LICENSE), covering the fine-tuned weights. - Apache License 2.0 — the Qwen3.8-27B base (
LICENSE-APACHE-2.0).
The Swift Open License limits use to non-commercial and research purposes (section 5, Commercial Use limitation). This derivative is redistributed under section 4, which permits Derivative Works provided the licence, attribution and
NOTICEtravel with them.NOTICEcarries a section 4(b) change notice describing exactly what was modified here. Where the two licences conflict, the more restrictive term governs.
Disclaimer
This model is provided for research purposes only, "AS IS", without warranty of any kind, express or implied. The author accepts no liability for any use of this model or any consequences arising from it. By downloading or using it, you accept sole responsibility for your use and for compliance with all applicable laws and regulations. Base model © Qwen (Apache-2.0); Swift contribution © UkisAI (Swift Open License v1.0); abliteration © the Heretic project; quantization via llm-compressor / compressed-tensors.
- Downloads last month
- 1,565
Model tree for akumaburn/Swift-1.5-Qwen3.8-27b-heretic-W8A8-DFlash2
Base model
Qwen/Qwen3.8-27B