Related models: all models

Swift-1.5-Qwen3.8-27B-Uncensored-BF16 · Swift-1.5-Qwen3.8-27B-Uncensored-FP8 · Swift-1.5-Qwen3.8-27B-Uncensored-FP8-NInfer · Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE · Swift-1.5-Qwen3.8-Flash-Next-Rank2-Abliteration-Patch · Swift-Qwen3.8-27B-FP8 · Swift-Qwen3.8-27B-Uncensored-BF16 · Swift-Qwen3.8-27B-Uncensored-FP8 · Swift-Qwen3.8-27B-Uncensored-NVFP4-LocalHessian-ActivationHeadroom-NInfer

Swift-1.5-Qwen3.8-27B-Uncensored-FP8

🔓 FP8 quantization of d0xin/Swift-1.5-Qwen3.8-27B-Uncensored-BF16.

The source BF16 checkpoint is an independent uncensored derivative of ukisai/Swift-1.5-Qwen3.8-27b, produced using rank-1 directional residual-stream ablation.

This release targets substantially lower model memory while retaining the Swift 1.5 architecture, multimodal support, tool calling, MTP weights and long-context configuration.

Initial capability and runtime evaluations are now available: 80.36% on the fixed 280-question MMLU-Pro subset, 85% IFEval prompt-strict / 90.18% instruction-strict, and 143.93 tok/s in the documented local NInfer agentic benchmark on an NVIDIA RTX PRO 6000 Blackwell 96 GB.

Additional refusal-behavior, dedicated multimodal and long-context evaluations remain separate work and will only be reported when validated specifically for this Swift 1.5 FP8 release.

Model lineage

Qwen/Qwen3.8-27B
        ↓
ukisai/Swift-1.5-Qwen3.8-27b
        ↓
rank-1 directional residual-stream ablation
        ↓
d0xin/Swift-1.5-Qwen3.8-27B-Uncensored-BF16
        ↓
FP8 quantization
        ↓
d0xin/Swift-1.5-Qwen3.8-27B-Uncensored-FP8

Model summary

  • FP8 checkpoint
  • approximately 29 GB
  • 18 safetensors shards
  • configured context: 262,144 tokens
  • source ablation layer: 38
  • source ablation rank: 1
  • modified residual writers: 131
  • vision tower unchanged by the ablation
  • MTP residual writers included
  • multimodal architecture retained
  • tool-calling architecture retained

FP8 conversion

The FP8 checkpoint was produced from the validated uncensored BF16 derivative.

Current artifact audit:

  • 18 safetensors shards
  • 1,606 tensors
  • 407 float8_e4m3fn tensors
  • 1,199 BF16 tensors
  • 407 scale tensors
  • safetensors integrity: PASS

The mixed dtypes are expected: tensors not selected for FP8 quantization remain in BF16 and quantized weights retain their associated scale tensors.

Ablation provenance

The BF16 source was created using:

  • source: ukisai/Swift-1.5-Qwen3.8-27b
  • ablation layer: 38
  • rank: 1
  • modified residual writers: 131
  • vision tower: unchanged
  • MTP residual writers: included

Transformation metadata is included in ABLITERATION.json.

Evaluation status

Initial capability and runtime evaluations are now available for this release.

The results below apply specifically to the Swift 1.5 FP8 derivative in this repository. Results from the earlier Swift-Qwen3.8 release or from the separate Flash-Next derivative are not mixed into this table.

Capability evaluation

Evaluation Result
MMLU-Pro 225 / 280 — 80.36%
IFEval — prompt strict 85%
IFEval — prompt loose 86%
IFEval — instruction strict 90.18%
IFEval — instruction loose 91.41%

The MMLU-Pro evaluation used a fixed 280-question subset consisting of 14 categories × 20 questions.

These results should be treated as reference measurements rather than comprehensive estimates of general model quality.

Agentic / reasoning runtime benchmark

A local xhigh reasoning run was also measured on the NInfer conversion derived from this FP8 checkpoint.

Test system

  • NVIDIA RTX PRO 6000 Blackwell 96 GB
  • NInfer runtime
  • concurrency: 1
  • FP8 KV cache
  • context / KV capacity: 65,536 tokens
  • prefill chunk: 1,024
  • DFlash2 speculative decoding
  • 5 draft tokens
  • LM-head draft enabled
  • thinking preserved

Measured run

Metric Result
Generation throughput 143.93 tok/s
Prefill throughput 1,625.62 tok/s
Completion tokens 21,764
Reasoning tokens 15,177
Final-answer tokens ~6,587
DFlash2 acceptance 39.9%

This is a runtime/workload-specific performance result, not a hardware-independent property of the checkpoint. Throughput will vary with inference engine, context length, concurrency, speculative-decoding configuration and hardware.

The corresponding NInfer release is:

d0xin/Swift-1.5-Qwen3.8-27B-Uncensored-FP8-NInfer

Validation coverage

Area Status
Checkpoint structure PASS
Safetensors integrity PASS
MMLU-Pro Completed
IFEval Completed
Agentic / reasoning performance Completed
Fixed 100-prompt refusal evaluation Not yet reported for Swift 1.5 FP8
Dedicated multimodal capability evaluation Not yet reported
Dedicated long-context retrieval evaluation Not yet reported

An early MATH-500 run is intentionally not reported because the evaluation adapter / answer-format handling for that run was invalid. Its resulting score must not be interpreted as a model capability result.

Further evaluations can be added to this model card without changing the released checkpoint.

Recommended generation settings

A good starting point based on the upstream Swift 1.5 configuration:

Parameter Value
reasoning_effort xhigh
temperature 1.0
top_p 0.95
top_k 20
min_p 0.0
presence_penalty 0.0
repetition_penalty 1.0

SGLang

python -m sglang.launch_server \
  --model-path d0xin/Swift-1.5-Qwen3.8-27B-Uncensored-FP8 \
  --served-model-name Swift-1.5-Qwen3.8-27B-Uncensored-FP8 \
  --trust-remote-code \
  --context-length 262144 \
  --reasoning-parser qwen3 \
  --tool-call-parser qwen3_coder \
  --port 30000

Adjust context length and memory parameters for the available hardware.

Tool calling

Use:

--tool-call-parser qwen3_coder

Multimodal support

The Qwen3.8 multimodal architecture and vision tower are retained.

The directional ablation used to create the BF16 source did not modify the vision tower.

Safety and responsible use

This is an uncensored / refusal-reduced derivative model.

The model has been intentionally modified to reduce refusal behavior. It may therefore generate content that the upstream model would normally refuse, restrict or handle more cautiously.

Outputs may be inaccurate, offensive, unsafe, unlawful or otherwise unsuitable for a particular use case.

Users are responsible for evaluating model outputs and ensuring compliance with applicable laws, licenses, regulations and platform requirements.

Upstream models

FP8 source:

d0xin/Swift-1.5-Qwen3.8-27B-Uncensored-BF16

Original Swift 1.5 model:

ukisai/Swift-1.5-Qwen3.8-27b

Base model:

Qwen/Qwen3.8-27B

License

These weights are distributed under the Swift Open License v1.0.

The underlying Qwen3.8 components remain subject to their applicable upstream license.

See the included LICENSE, LICENSE-APACHE-2.0, and NOTICE files for the applicable terms and attribution requirements.

Attribution

  • Base model: Qwen/Qwen3.8-27B
  • Swift post-training: UkisAI
  • Swift 1.5: ukisai/Swift-1.5-Qwen3.8-27b
  • Directional ablation: d0xin
  • FP8 quantization and release packaging: d0xin

Citation

@misc{swift-1.5-qwen3.8-27b,
  title  = {Swift 1.5 Qwen3.8-27B},
  author = {UkisAI},
  year   = {2026},
  url    = {https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b}
}
Downloads last month
498
Safetensors
Model size
28B params
Tensor type
BF16
·
F8_E4M3
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for d0xin/Swift-1.5-Qwen3.8-27B-Uncensored-FP8

Base model

Qwen/Qwen3.8-27B
Quantized
(4)
this model
Finetunes
1 model