Related models: all models
Swift-1.5-Qwen3.8-27B-Uncensored-BF16 · Swift-1.5-Qwen3.8-27B-Uncensored-FP8 · Swift-1.5-Qwen3.8-27B-Uncensored-FP8-NInfer · Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE · Swift-1.5-Qwen3.8-Flash-Next-Rank2-Abliteration-Patch · Swift-Qwen3.8-27B-FP8 · Swift-Qwen3.8-27B-Uncensored-BF16 · Swift-Qwen3.8-27B-Uncensored-FP8 · Swift-Qwen3.8-27B-Uncensored-NVFP4-LocalHessian-ActivationHeadroom-NInfer
Swift-1.5 Qwen3.8 Flash-Next — Rank-2 Abliteration Patch
Reproducible Rank-2 abliteration patch for:
d0xin/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE
This repo contains no model weights. It provides the validated rank-2 projection basis, transformation plan, baker, reference manifest and one-command patch script.
Transformation
- basis:
{Orca, Swift⊥} - rank: 2
- alpha: 1.0
- BF16 targets: 100
- NVFP4 expert weights: 24,576
- companion scale tensors: 49,152
- modified shards: 49
- modified payload: 78.227 GiB
- 10 FP8 PLE shards: unchanged
Ablation result and behavioral tests
The purpose of this transform is to suppress the learned refusal direction while preserving the underlying model capabilities.
Refusal behavior
The validated rank-2 reference transform produced:
| Evaluation | Result |
|---|---|
| Primary validation set | 80 / 80 DIRECT (100%) |
| Held-out JBB non-AdvBench set | 79 / 80 DIRECT (98.75%) |
| Held-out refusals | 1 / 80 (1.25%) |
The held-out evaluation used jbb-non-advbench-heldout.csv with:
- 80 previously held-out prompts
- temperature:
0 - max output:
192tokens - concurrency:
4 - deterministic refusal-marker classifier
- output classified as
REFUSEwhen a refusal marker was detected, otherwiseDIRECT
The held-out set was intentionally kept separate from the prompts used while developing the projection.
Performance
Reference generation-speed runs with the rank-2 transform:
| Run | Throughput |
|---|---|
| 1 | 141.75 tok/s |
| 2 | 138.66 tok/s |
| 3 | 154.47 tok/s |
| Median | 141.75 tok/s |
These measurements were obtained on the RTX PRO 6000 96 GB reference deployment using Pennyroyal / SGLang v2.5.3.
Functional regression checks
The validated runtime configuration also passed:
- normal chat-completion inference
- native NEXTN / MTP speculative decoding
- structured startup warmup
- tool calling with correctly formed function arguments
- end-to-end routing through
hybrid_auto - 262,144-token configured context
- NVMe-backed FP8 PLE operation
No agent/tool-calling regression was observed in these smoke tests.
Important measurement note
The behavioral scores above were measured with the validated runtime application of the same rank-2 {Orca, Swift⊥} projection represented by this patch kit.
The published offline baker was then validated independently against the baked reference checkpoint:
- all 49 modified shard SHA256 hashes match the reference build;
- transformation totals match exactly;
- a tensor-level differential smoke test changed all 1,538 expected tensors and zero unrelated tensors.
Therefore the patch kit is designed to reproduce the validated transformation exactly. The 79/80 behavioral benchmark itself was not rerun as a separate full benchmark after packaging this HF patch-kit release.
Base checkpoint
hf download d0xin/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE \
--local-dir ./Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE
Validated fingerprints:
config.json:30a6f84424ff7481a9f16975b1e350e1563dd66c20e72a23b47cde6047c8cd71model.safetensors.index.json:c80a68f96121ce1e38b5cd9e62b7c4e23906714dbfb3ea914b0294c6c43580e7
Apply
Requirements: Linux, Docker, NVIDIA Container Toolkit, NVIDIA GPU, about 80 GiB additional free disk space. BASE and OUTPUT should be on the same filesystem so unchanged PLE shards can be hardlinked.
./apply_patch.sh \
/path/to/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE \
/path/to/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE-Rank2-Abliterated
Successful completion:
REFERENCE_MANIFEST=PASS
bf16=100
experts=24576
side=49152
PATCH_APPLY=PASS
RTX PRO 6000 96 GB / Pennyroyal
Reference runtime:
ghcr.io/jpezzulli/sglang-rtxpro6000:v2.5.3
After baking, prepare the Pennyroyal NVMe PLE derivative from the baked checkpoint and run with:
TARGET_MODEL=<full baked checkpoint>
PENNY_PLE_BACKEND=nvme
PENNY_PLE_NVME_MODEL=<prepared NVMe PLE derivative>
The large PLE table is streamed from SSD/NVMe rather than kept resident in GPU memory.
Pennyroyal: https://github.com/jpezzulli/sglang-rtxpro6000
Reference validation
The reference build passed:
- 59/59 safetensors shards
- 296,475/296,475 indexed tensors
- 0 missing tensors
- 0 extra tensors
- 0 duplicate tensors
- 0 wrong-shard mappings
- 49/49 modified shard SHA256 checks
Differential smoke audit:
expected_touched = 1538
touched_changed = 1538
touched_same = 0
unexpected_changed = 0
Reproducibility hashes
- direction:
8eb11e23dcbea3cbc040f5077856b19ea475d5af81d1a62f789ebe96d5130b2c - plan:
22ab558cf8708bf1f6a57d475ce2bc0c9ce7b5a9a54e5f756be70e28f30a7f50 - Stage-B implementation:
5138983ee95bd80b2645ce8f37ff20bab0c023376f6f68ee0c567842a6ebfbcf - baker:
e777252cdb7fa445c4d19cd316f3218d736de9af2a31e9b4ba56b9c6322f8e8c
Included
direction_rank2_orca_swift.ptplan.jsonflashnext_abliteration_stage_b_rank2.pybake_flashnext_rank2.pyapply_patch.shreference_BAKE_MANIFEST.jsonlBASE_FINGERPRINTS.jsonSHA256SUMS
Safety
This is an abliteration / uncensoring experiment and intentionally reduces refusal behavior. Deployment operators are responsible for appropriate application-level safety controls.
Model tree for d0xin/Swift-1.5-Qwen3.8-Flash-Next-Rank2-Abliteration-Patch
Base model
Qwen/Qwen3.8-Flash-Next