Swift-1.5-Qwen3.8-27B — Heretic Abliterated (BF16)
Research artifact. Safety alignment has been deliberately removed. This model will attempt to comply with harmful, dangerous, illegal, and unethical requests that the source model refuses, with no content moderation. See Safety. Provided for research purposes only, with no warranty and no liability accepted by the author — see the disclaimer at the bottom.
An abliterated build of ukisai/Swift-1.5-Qwen3.8-27b, UkisAI's
reasoning-efficient derivative of Qwen3.8-27B.
Refusal directions are removed by directional ablation; nothing else is retrained.
- Precision: BF16, unchanged from the source.
- Untouched: the vision tower, the MTP speculative-decoding head (bitwise
identical to the source's),
lm_head,embed_tokens, all norms, and every GatedDeltaNet recurrent gate. - Size 52G, 1,199 tensors (333 vision, 15 MTP).
Abliteration
Produced with Heretic v2.0.0.dev0: 260 TPE trials,
seed 42, directional ablation on attn.o_proj (16 modules), attn.out_proj (48) and
mlp.down_proj (64). Baseline refusal score before abliteration: 98/100.
Trial 140 was selected — Pareto index 1, not index 0. Index 0 (trial 161) scored 22/100 keywords at KL 0.1091; trial 140 scores 23/100 at KL 0.0602. One extra soft keyword hit in exchange for 45% less divergence from the source was the better trade, and the refusal breakdown above confirms the cost was nil: both are 0/100 on genuine refusals.
Independent verification on 100 harmless prompts (full-vocabulary first-token KL,
outside Heretic) measures KL(source ‖ this) = 0.0616, agreeing with
Heretic's own 0.0602 to within 2%.
This is a thinking model: the chat template emits <think> as part of the generation
prompt, so scoring used the response prefix \n</think>\n\n to close the block —
without it every scored token is reasoning text and the refusal metric saturates.
Variants
All four builds derive from the same trial-140 abliteration. Measured on identical data and hardware (see Benchmarks).
| build | size | first-token KL vs source | quant-only KL | accept. length | best for |
|---|---|---|---|---|---|
| BF16 abliterated | 52G | 0.0616 | — | n/a | research, re-quantization |
| W8A8 + QuaRot/SmoothQuant | 30G | 0.0843 | 0.0264 | 2.51 (MTP-3) | highest fidelity; batch serving |
| W8A8 unrotated | 29G | 0.1330 | 0.0791 | 4.34 (DFlash2) | DFlash2 speculation at c ≤ 8 |
| W4A16 unrotated | 18G | 0.2109 | 0.1182 | 3.61 (DFlash2) | smallest; single-user decode |
quant-only KL isolates the quantization from the abliteration: KL(abliterated ‖ build).
The rotation is worth 3.0× on this metric (0.0264 vs 0.0791 unrotated),
and 4-bit costs 4.5× the rotated 8-bit build.
Match speculative depth to concurrency. DFlash2 (k=7) gives the best decode at low concurrency, but its prefill collapses when the GPU saturates: on the unrotated W8A8, prefill falls from 7,329 tok/s at c=8 to 2,829 tok/s at c=32, and TTFT p50 rises to 12.6 s. The rotated W8A8 with MTP-3 instead climbs to 19,209 tok/s at c=32. Use DFlash2 for c ≤ 8; prefer the rotated W8A8 with MTP for c ≥ 16.
Benchmarks
Measured on 2× NVIDIA CMP 170HX (64 GB, sm_80), TP=1 (single card per server),
vLLM 0.27.1, fp8_e4m3 KV cache, async scheduling, --max-model-len 32768.
Synthetic input 1,024 tokens / output 256, --ignore-eos; acceptance measured
separately on 88 code tasks, greedy, concurrency 1.
vLLM 0.27.1 was used deliberately: on this Ampere hardware, --kv-cache-dtype fp8_e4m3
makes vLLM 0.28.0 fall back to FlashInfer (FlashAttention has no fp8-KV path below
Hopper), which faults with an illegal memory access under speculative decoding.
| build | c | prefill tok/s | decode tok/s | TTFT p50 | TPOT p50 | accept. len |
|---|---|---|---|---|---|---|
| W8A8 rotated (MTP-3) | 1 | 3,432 | 57.3 | 272 ms | 13.12 ms | 2.06 |
| 8 | 8,720 | 290.7 | 350 ms | 18.87 ms | 2.61 | |
| 32 | 19,209 | 516.8 | 643 ms | 46.29 ms | 2.41 | |
| W8A8 unrot (DFlash2 k=7) | 1 | 2,882 | 87.5 | 337 ms | 8.85 ms | 3.12 |
| 8 | 7,329 | 321.1 | 936 ms | 17.17 ms | 3.38 | |
| 32 | 2,829 | 410.8 | 12551 ms | 21.99 ms | 3.66 | |
| W4A16 (DFlash2 k=7) | 1 | 1,926 | 119.3 | 514 ms | 5.88 ms | 3.63 |
| 8 | 4,983 | 254.5 | 1240 ms | 21.08 ms | 3.39 | |
| 32 | 1,901 | 280.2 | 18356 ms | 32.48 ms | 3.72 |
Acceptance on 88 code tasks, greedy, c=1: W8A8 rotated (MTP-3) 50.3% / 2.51,
W8A8 unrotated (DFlash2 k=7) 47.7% / 4.34,
W4A16 (DFlash2 k=7) 37.3% / 3.61.
Acceptance length (mean tokens accepted per verify step) is what sets the speedup;
the rate alone is not comparable across different num_speculative_tokens.
The source card's task benchmarks (GPQA-Diamond, IFBench, AIME 2026, LiveCodeBench, Terminal-Bench) were not re-run here — single-GPU budget. The published values on the source card stand as reported, and nothing in this build is claimed to preserve them. The KL figures above are the fidelity evidence actually measured.
Usage
from transformers import AutoModelForImageTextToText, AutoProcessor
m = AutoModelForImageTextToText.from_pretrained(
"akumaburn/Swift-1.5-Qwen3.8-27b-heretic", dtype="bfloat16", device_map="auto")
Vision is intact — feed images at full resolution; small images yield few visual tokens and degrade OCR noticeably.
Safety
Refusal behaviour has been deliberately removed. Measured on 100 harmful-instruction prompts, with the keyword classes separated:
| count | |
|---|---|
| Hard refusals (genuine refusal language) | 0 / 100 |
| Soft caveats only (complied, but said "illegal"/"harmful"/…) | 24 / 100 |
| True refusal rate | 0 / 100 |
The residual 24/100 keyword hits are compliant answers that add a caveat, not refusals — a measurement floor of the KeywordRate metric rather than surviving censorship.
This model produces content the source declines, including dangerous, illegal, or unethical material, with no moderation. Intended for interpretability/safety research, red-teaming, and evaluation by people who understand and accept those risks. Do not deploy it where it can reach people who have not consented to unfiltered output. You are responsible for your use and for compliance with all applicable laws.
Licence
Dual stack (see NOTICE and the two licence files shipped with this model):
- Swift Open License v1.0 — UkisAI's contribution (
LICENSE), covering the fine-tuned weights. - Apache License 2.0 — the Qwen3.8-27B base (
LICENSE-APACHE-2.0).
The Swift Open License limits use to non-commercial and research purposes (section 5, Commercial Use limitation). This derivative is redistributed under section 4, which permits Derivative Works provided the licence, attribution and
NOTICEtravel with them.NOTICEcarries a section 4(b) change notice describing exactly what was modified here. Where the two licences conflict, the more restrictive term governs.
Disclaimer
This model is provided for research purposes only, "AS IS", without warranty of any kind, express or implied. The author accepts no liability for any use of this model or any consequences arising from it. By downloading or using it, you accept sole responsibility for your use and for compliance with all applicable laws and regulations. Base model © Qwen (Apache-2.0); Swift contribution © UkisAI (Swift Open License v1.0); abliteration © the Heretic project; quantization via llm-compressor / compressed-tensors.
- Downloads last month
- 77