Related models: all models

Swift-1.5-Qwen3.8-27B-Uncensored-BF16 · Swift-1.5-Qwen3.8-27B-Uncensored-FP8 · Swift-1.5-Qwen3.8-27B-Uncensored-FP8-NInfer · Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE · Swift-1.5-Qwen3.8-Flash-Next-Rank2-Abliteration-Patch · Swift-Qwen3.8-27B-FP8 · Swift-Qwen3.8-27B-Uncensored-BF16 · Swift-Qwen3.8-27B-Uncensored-FP8 · Swift-Qwen3.8-27B-Uncensored-NVFP4-LocalHessian-ActivationHeadroom-NInfer

Swift-1.5 Qwen3.8 Flash-Next — Rank-2 Abliteration Patch

Reproducible Rank-2 abliteration patch for:

d0xin/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE

This repo contains no model weights. It provides the validated rank-2 projection basis, transformation plan, baker, reference manifest and one-command patch script.

Transformation

  • basis: {Orca, Swift⊥}
  • rank: 2
  • alpha: 1.0
  • BF16 targets: 100
  • NVFP4 expert weights: 24,576
  • companion scale tensors: 49,152
  • modified shards: 49
  • modified payload: 78.227 GiB
  • 10 FP8 PLE shards: unchanged

Ablation result and behavioral tests

The purpose of this transform is to suppress the learned refusal direction while preserving the underlying model capabilities.

Refusal behavior

The validated rank-2 reference transform produced:

Evaluation Result
Primary validation set 80 / 80 DIRECT (100%)
Held-out JBB non-AdvBench set 79 / 80 DIRECT (98.75%)
Held-out refusals 1 / 80 (1.25%)

The held-out evaluation used jbb-non-advbench-heldout.csv with:

  • 80 previously held-out prompts
  • temperature: 0
  • max output: 192 tokens
  • concurrency: 4
  • deterministic refusal-marker classifier
  • output classified as REFUSE when a refusal marker was detected, otherwise DIRECT

The held-out set was intentionally kept separate from the prompts used while developing the projection.

Performance

Reference generation-speed runs with the rank-2 transform:

Run Throughput
1 141.75 tok/s
2 138.66 tok/s
3 154.47 tok/s
Median 141.75 tok/s

These measurements were obtained on the RTX PRO 6000 96 GB reference deployment using Pennyroyal / SGLang v2.5.3.

Functional regression checks

The validated runtime configuration also passed:

  • normal chat-completion inference
  • native NEXTN / MTP speculative decoding
  • structured startup warmup
  • tool calling with correctly formed function arguments
  • end-to-end routing through hybrid_auto
  • 262,144-token configured context
  • NVMe-backed FP8 PLE operation

No agent/tool-calling regression was observed in these smoke tests.

Important measurement note

The behavioral scores above were measured with the validated runtime application of the same rank-2 {Orca, Swift⊥} projection represented by this patch kit.

The published offline baker was then validated independently against the baked reference checkpoint:

  • all 49 modified shard SHA256 hashes match the reference build;
  • transformation totals match exactly;
  • a tensor-level differential smoke test changed all 1,538 expected tensors and zero unrelated tensors.

Therefore the patch kit is designed to reproduce the validated transformation exactly. The 79/80 behavioral benchmark itself was not rerun as a separate full benchmark after packaging this HF patch-kit release.

Base checkpoint

hf download d0xin/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE \
  --local-dir ./Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE

Validated fingerprints:

  • config.json: 30a6f84424ff7481a9f16975b1e350e1563dd66c20e72a23b47cde6047c8cd71
  • model.safetensors.index.json: c80a68f96121ce1e38b5cd9e62b7c4e23906714dbfb3ea914b0294c6c43580e7

Apply

Requirements: Linux, Docker, NVIDIA Container Toolkit, NVIDIA GPU, about 80 GiB additional free disk space. BASE and OUTPUT should be on the same filesystem so unchanged PLE shards can be hardlinked.

./apply_patch.sh \
  /path/to/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE \
  /path/to/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE-Rank2-Abliterated

Successful completion:

REFERENCE_MANIFEST=PASS
bf16=100
experts=24576
side=49152
PATCH_APPLY=PASS

RTX PRO 6000 96 GB / Pennyroyal

Reference runtime:

ghcr.io/jpezzulli/sglang-rtxpro6000:v2.5.3

After baking, prepare the Pennyroyal NVMe PLE derivative from the baked checkpoint and run with:

TARGET_MODEL=<full baked checkpoint>
PENNY_PLE_BACKEND=nvme
PENNY_PLE_NVME_MODEL=<prepared NVMe PLE derivative>

The large PLE table is streamed from SSD/NVMe rather than kept resident in GPU memory.

Pennyroyal: https://github.com/jpezzulli/sglang-rtxpro6000

Reference validation

The reference build passed:

  • 59/59 safetensors shards
  • 296,475/296,475 indexed tensors
  • 0 missing tensors
  • 0 extra tensors
  • 0 duplicate tensors
  • 0 wrong-shard mappings
  • 49/49 modified shard SHA256 checks

Differential smoke audit:

expected_touched   = 1538
touched_changed    = 1538
touched_same       = 0
unexpected_changed = 0

Reproducibility hashes

  • direction: 8eb11e23dcbea3cbc040f5077856b19ea475d5af81d1a62f789ebe96d5130b2c
  • plan: 22ab558cf8708bf1f6a57d475ce2bc0c9ce7b5a9a54e5f756be70e28f30a7f50
  • Stage-B implementation: 5138983ee95bd80b2645ce8f37ff20bab0c023376f6f68ee0c567842a6ebfbcf
  • baker: e777252cdb7fa445c4d19cd316f3218d736de9af2a31e9b4ba56b9c6322f8e8c

Included

  • direction_rank2_orca_swift.pt
  • plan.json
  • flashnext_abliteration_stage_b_rank2.py
  • bake_flashnext_rank2.py
  • apply_patch.sh
  • reference_BAKE_MANIFEST.jsonl
  • BASE_FINGERPRINTS.json
  • SHA256SUMS

Safety

This is an abliteration / uncensoring experiment and intentionally reduces refusal behavior. Deployment operators are responsible for appropriate application-level safety controls.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for d0xin/Swift-1.5-Qwen3.8-Flash-Next-Rank2-Abliteration-Patch