Swift-1.5-Qwen3.8-27B — Heretic Abliterated · GPTQ W8A8 INT8 (unrotated, DFlash2)

Research artifact. Safety alignment has been deliberately removed. This model will attempt to comply with harmful, dangerous, illegal, and unethical requests that the source model refuses, with no content moderation. See Safety. Provided for research purposes only, with no warranty and no liability accepted by the author — see the disclaimer at the bottom.

An INT8 W8A8 build of akumaburn/Swift-1.5-Qwen3.8-27b-heretic deliberately produced without the QuaRot rotation, so the DFlash2 drafter works. Size 29G.

  • Weights: INT8, per-channel symmetric (GPTQ).
  • Activations: INT8, per-token dynamic (int-quantized).
  • Unrotated: norms are preserved (not folded to zero), which is what DFlash2's fc projection requires.
  • Left unquantized: MTP head, lm_head, embed_tokens, vision tower, GatedDeltaNet gates, all norms.

Fidelity — the cost of staying unrotated

value
KL(abliterated ‖ this) — quantization only 0.0791
KL(source ‖ this) 0.1330

This is 3.0× the divergence of the rotated sibling (0.0264) at the same bit-width and nearly the same size. That is the price of DFlash2 compatibility, and it buys the highest acceptance length of the set (4.34 vs 2.51 for MTP-3).

Abliteration

Produced with Heretic v2.0.0.dev0: 260 TPE trials, seed 42, directional ablation on attn.o_proj (16 modules), attn.out_proj (48) and mlp.down_proj (64). Baseline refusal score before abliteration: 98/100.

Trial 140 was selected — Pareto index 1, not index 0. Index 0 (trial 161) scored 22/100 keywords at KL 0.1091; trial 140 scores 23/100 at KL 0.0602. One extra soft keyword hit in exchange for 45% less divergence from the source was the better trade, and the refusal breakdown above confirms the cost was nil: both are 0/100 on genuine refusals.

Independent verification on 100 harmless prompts (full-vocabulary first-token KL, outside Heretic) measures KL(source ‖ this) = 0.0616, agreeing with Heretic's own 0.0602 to within 2%.

This is a thinking model: the chat template emits <think> as part of the generation prompt, so scoring used the response prefix \n</think>\n\n to close the block — without it every scored token is reasoning text and the refusal metric saturates.

Variants

All four builds derive from the same trial-140 abliteration. Measured on identical data and hardware (see Benchmarks).

build size first-token KL vs source quant-only KL accept. length best for
BF16 abliterated 52G 0.0616 — n/a research, re-quantization
W8A8 + QuaRot/SmoothQuant 30G 0.0843 0.0264 2.51 (MTP-3) highest fidelity; batch serving
W8A8 unrotated 29G 0.1330 0.0791 4.34 (DFlash2) DFlash2 speculation at c ≤ 8
W4A16 unrotated 18G 0.2109 0.1182 3.61 (DFlash2) smallest; single-user decode

quant-only KL isolates the quantization from the abliteration: KL(abliterated ‖ build). The rotation is worth 3.0× on this metric (0.0264 vs 0.0791 unrotated), and 4-bit costs 4.5× the rotated 8-bit build.

Match speculative depth to concurrency. DFlash2 (k=7) gives the best decode at low concurrency, but its prefill collapses when the GPU saturates: on the unrotated W8A8, prefill falls from 7,329 tok/s at c=8 to 2,829 tok/s at c=32, and TTFT p50 rises to 12.6 s. The rotated W8A8 with MTP-3 instead climbs to 19,209 tok/s at c=32. Use DFlash2 for c ≤ 8; prefer the rotated W8A8 with MTP for c ≥ 16.

Benchmarks

Measured on 2× NVIDIA CMP 170HX (64 GB, sm_80), TP=1 (single card per server), vLLM 0.27.1, fp8_e4m3 KV cache, async scheduling, --max-model-len 32768. Synthetic input 1,024 tokens / output 256, --ignore-eos; acceptance measured separately on 88 code tasks, greedy, concurrency 1.

vLLM 0.27.1 was used deliberately: on this Ampere hardware, --kv-cache-dtype fp8_e4m3 makes vLLM 0.28.0 fall back to FlashInfer (FlashAttention has no fp8-KV path below Hopper), which faults with an illegal memory access under speculative decoding.

build c prefill tok/s decode tok/s TTFT p50 TPOT p50 accept. len
W8A8 rotated (MTP-3) 1 3,432 57.3 272 ms 13.12 ms 2.06
8 8,720 290.7 350 ms 18.87 ms 2.61
32 19,209 516.8 643 ms 46.29 ms 2.41
W8A8 unrot (DFlash2 k=7) 1 2,882 87.5 337 ms 8.85 ms 3.12
8 7,329 321.1 936 ms 17.17 ms 3.38
32 2,829 410.8 12551 ms 21.99 ms 3.66
W4A16 (DFlash2 k=7) 1 1,926 119.3 514 ms 5.88 ms 3.63
8 4,983 254.5 1240 ms 21.08 ms 3.39
32 1,901 280.2 18356 ms 32.48 ms 3.72

Acceptance on 88 code tasks, greedy, c=1: W8A8 rotated (MTP-3) 50.3% / 2.51, W8A8 unrotated (DFlash2 k=7) 47.7% / 4.34, W4A16 (DFlash2 k=7) 37.3% / 3.61. Acceptance length (mean tokens accepted per verify step) is what sets the speedup; the rate alone is not comparable across different num_speculative_tokens.

The source card's task benchmarks (GPQA-Diamond, IFBench, AIME 2026, LiveCodeBench, Terminal-Bench) were not re-run here — single-GPU budget. The published values on the source card stand as reported, and nothing in this build is claimed to preserve them. The KL figures above are the fidelity evidence actually measured.

Serving (vLLM, with DFlash2)

vllm serve akumaburn/Swift-1.5-Qwen3.8-27b-heretic-W8A8-DFlash2 \
  --tensor-parallel-size 1 --max-model-len 32768 \
  --kv-cache-dtype fp8_e4m3 --async-scheduling \
  --speculative-config '{"method":"dflash","model":"<DFlash2 drafter>","num_speculative_tokens":7}'

Drop the --speculative-config for concurrency ≥ 16; see the tip above.

Safety

Refusal behaviour has been deliberately removed. Measured on 100 harmful-instruction prompts, with the keyword classes separated:

count
Hard refusals (genuine refusal language) 0 / 100
Soft caveats only (complied, but said "illegal"/"harmful"/…) 24 / 100
True refusal rate 0 / 100

The residual 24/100 keyword hits are compliant answers that add a caveat, not refusals — a measurement floor of the KeywordRate metric rather than surviving censorship.

This model produces content the source declines, including dangerous, illegal, or unethical material, with no moderation. Intended for interpretability/safety research, red-teaming, and evaluation by people who understand and accept those risks. Do not deploy it where it can reach people who have not consented to unfiltered output. You are responsible for your use and for compliance with all applicable laws.

Licence

Dual stack (see NOTICE and the two licence files shipped with this model):

  • Swift Open License v1.0 — UkisAI's contribution (LICENSE), covering the fine-tuned weights.
  • Apache License 2.0 — the Qwen3.8-27B base (LICENSE-APACHE-2.0).

The Swift Open License limits use to non-commercial and research purposes (section 5, Commercial Use limitation). This derivative is redistributed under section 4, which permits Derivative Works provided the licence, attribution and NOTICE travel with them. NOTICE carries a section 4(b) change notice describing exactly what was modified here. Where the two licences conflict, the more restrictive term governs.

Disclaimer

This model is provided for research purposes only, "AS IS", without warranty of any kind, express or implied. The author accepts no liability for any use of this model or any consequences arising from it. By downloading or using it, you accept sole responsibility for your use and for compliance with all applicable laws and regulations. Base model © Qwen (Apache-2.0); Swift contribution © UkisAI (Swift Open License v1.0); abliteration © the Heretic project; quantization via llm-compressor / compressed-tensors.

Downloads last month
1,565
Safetensors
Model size
27B params
Tensor type
BF16
·
I8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for akumaburn/Swift-1.5-Qwen3.8-27b-heretic-W8A8-DFlash2

Base model

Qwen/Qwen3.8-27B
Quantized
(5)
this model