Swift-1.5-Qwen3.8-27B — Heretic Abliterated (BF16)

Research artifact. Safety alignment has been deliberately removed. This model will attempt to comply with harmful, dangerous, illegal, and unethical requests that the source model refuses, with no content moderation. See Safety. Provided for research purposes only, with no warranty and no liability accepted by the author — see the disclaimer at the bottom.

An abliterated build of ukisai/Swift-1.5-Qwen3.8-27b, UkisAI's reasoning-efficient derivative of Qwen3.8-27B. Refusal directions are removed by directional ablation; nothing else is retrained.

  • Precision: BF16, unchanged from the source.
  • Untouched: the vision tower, the MTP speculative-decoding head (bitwise identical to the source's), lm_head, embed_tokens, all norms, and every GatedDeltaNet recurrent gate.
  • Size 52G, 1,199 tensors (333 vision, 15 MTP).

Abliteration

Produced with Heretic v2.0.0.dev0: 260 TPE trials, seed 42, directional ablation on attn.o_proj (16 modules), attn.out_proj (48) and mlp.down_proj (64). Baseline refusal score before abliteration: 98/100.

Trial 140 was selected — Pareto index 1, not index 0. Index 0 (trial 161) scored 22/100 keywords at KL 0.1091; trial 140 scores 23/100 at KL 0.0602. One extra soft keyword hit in exchange for 45% less divergence from the source was the better trade, and the refusal breakdown above confirms the cost was nil: both are 0/100 on genuine refusals.

Independent verification on 100 harmless prompts (full-vocabulary first-token KL, outside Heretic) measures KL(source ‖ this) = 0.0616, agreeing with Heretic's own 0.0602 to within 2%.

This is a thinking model: the chat template emits <think> as part of the generation prompt, so scoring used the response prefix \n</think>\n\n to close the block — without it every scored token is reasoning text and the refusal metric saturates.

Variants

All four builds derive from the same trial-140 abliteration. Measured on identical data and hardware (see Benchmarks).

build size first-token KL vs source quant-only KL accept. length best for
BF16 abliterated 52G 0.0616 — n/a research, re-quantization
W8A8 + QuaRot/SmoothQuant 30G 0.0843 0.0264 2.51 (MTP-3) highest fidelity; batch serving
W8A8 unrotated 29G 0.1330 0.0791 4.34 (DFlash2) DFlash2 speculation at c ≤ 8
W4A16 unrotated 18G 0.2109 0.1182 3.61 (DFlash2) smallest; single-user decode

quant-only KL isolates the quantization from the abliteration: KL(abliterated ‖ build). The rotation is worth 3.0× on this metric (0.0264 vs 0.0791 unrotated), and 4-bit costs 4.5× the rotated 8-bit build.

Match speculative depth to concurrency. DFlash2 (k=7) gives the best decode at low concurrency, but its prefill collapses when the GPU saturates: on the unrotated W8A8, prefill falls from 7,329 tok/s at c=8 to 2,829 tok/s at c=32, and TTFT p50 rises to 12.6 s. The rotated W8A8 with MTP-3 instead climbs to 19,209 tok/s at c=32. Use DFlash2 for c ≤ 8; prefer the rotated W8A8 with MTP for c ≥ 16.

Benchmarks

Measured on 2× NVIDIA CMP 170HX (64 GB, sm_80), TP=1 (single card per server), vLLM 0.27.1, fp8_e4m3 KV cache, async scheduling, --max-model-len 32768. Synthetic input 1,024 tokens / output 256, --ignore-eos; acceptance measured separately on 88 code tasks, greedy, concurrency 1.

vLLM 0.27.1 was used deliberately: on this Ampere hardware, --kv-cache-dtype fp8_e4m3 makes vLLM 0.28.0 fall back to FlashInfer (FlashAttention has no fp8-KV path below Hopper), which faults with an illegal memory access under speculative decoding.

build c prefill tok/s decode tok/s TTFT p50 TPOT p50 accept. len
W8A8 rotated (MTP-3) 1 3,432 57.3 272 ms 13.12 ms 2.06
8 8,720 290.7 350 ms 18.87 ms 2.61
32 19,209 516.8 643 ms 46.29 ms 2.41
W8A8 unrot (DFlash2 k=7) 1 2,882 87.5 337 ms 8.85 ms 3.12
8 7,329 321.1 936 ms 17.17 ms 3.38
32 2,829 410.8 12551 ms 21.99 ms 3.66
W4A16 (DFlash2 k=7) 1 1,926 119.3 514 ms 5.88 ms 3.63
8 4,983 254.5 1240 ms 21.08 ms 3.39
32 1,901 280.2 18356 ms 32.48 ms 3.72

Acceptance on 88 code tasks, greedy, c=1: W8A8 rotated (MTP-3) 50.3% / 2.51, W8A8 unrotated (DFlash2 k=7) 47.7% / 4.34, W4A16 (DFlash2 k=7) 37.3% / 3.61. Acceptance length (mean tokens accepted per verify step) is what sets the speedup; the rate alone is not comparable across different num_speculative_tokens.

The source card's task benchmarks (GPQA-Diamond, IFBench, AIME 2026, LiveCodeBench, Terminal-Bench) were not re-run here — single-GPU budget. The published values on the source card stand as reported, and nothing in this build is claimed to preserve them. The KL figures above are the fidelity evidence actually measured.

Usage

from transformers import AutoModelForImageTextToText, AutoProcessor
m = AutoModelForImageTextToText.from_pretrained(
    "akumaburn/Swift-1.5-Qwen3.8-27b-heretic", dtype="bfloat16", device_map="auto")

Vision is intact — feed images at full resolution; small images yield few visual tokens and degrade OCR noticeably.

Safety

Refusal behaviour has been deliberately removed. Measured on 100 harmful-instruction prompts, with the keyword classes separated:

count
Hard refusals (genuine refusal language) 0 / 100
Soft caveats only (complied, but said "illegal"/"harmful"/…) 24 / 100
True refusal rate 0 / 100

The residual 24/100 keyword hits are compliant answers that add a caveat, not refusals — a measurement floor of the KeywordRate metric rather than surviving censorship.

This model produces content the source declines, including dangerous, illegal, or unethical material, with no moderation. Intended for interpretability/safety research, red-teaming, and evaluation by people who understand and accept those risks. Do not deploy it where it can reach people who have not consented to unfiltered output. You are responsible for your use and for compliance with all applicable laws.

Licence

Dual stack (see NOTICE and the two licence files shipped with this model):

  • Swift Open License v1.0 — UkisAI's contribution (LICENSE), covering the fine-tuned weights.
  • Apache License 2.0 — the Qwen3.8-27B base (LICENSE-APACHE-2.0).

The Swift Open License limits use to non-commercial and research purposes (section 5, Commercial Use limitation). This derivative is redistributed under section 4, which permits Derivative Works provided the licence, attribution and NOTICE travel with them. NOTICE carries a section 4(b) change notice describing exactly what was modified here. Where the two licences conflict, the more restrictive term governs.

Disclaimer

This model is provided for research purposes only, "AS IS", without warranty of any kind, express or implied. The author accepts no liability for any use of this model or any consequences arising from it. By downloading or using it, you accept sole responsibility for your use and for compliance with all applicable laws and regulations. Base model © Qwen (Apache-2.0); Swift contribution © UkisAI (Swift Open License v1.0); abliteration © the Heretic project; quantization via llm-compressor / compressed-tensors.

Downloads last month
77
Safetensors
Model size
27B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for akumaburn/Swift-1.5-Qwen3.8-27b-heretic

Base model

Qwen/Qwen3.8-27B
Finetuned
(6)
this model
Quantizations
5 models