Instructions to use d0xin/Swift-1.5-Qwen3.8-27B-Uncensored-FP8 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use d0xin/Swift-1.5-Qwen3.8-27B-Uncensored-FP8 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="d0xin/Swift-1.5-Qwen3.8-27B-Uncensored-FP8") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("d0xin/Swift-1.5-Qwen3.8-27B-Uncensored-FP8") model = AutoModelForMultimodalLM.from_pretrained("d0xin/Swift-1.5-Qwen3.8-27B-Uncensored-FP8", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use d0xin/Swift-1.5-Qwen3.8-27B-Uncensored-FP8 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "d0xin/Swift-1.5-Qwen3.8-27B-Uncensored-FP8" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "d0xin/Swift-1.5-Qwen3.8-27B-Uncensored-FP8", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/d0xin/Swift-1.5-Qwen3.8-27B-Uncensored-FP8
- SGLang
How to use d0xin/Swift-1.5-Qwen3.8-27B-Uncensored-FP8 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "d0xin/Swift-1.5-Qwen3.8-27B-Uncensored-FP8" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "d0xin/Swift-1.5-Qwen3.8-27B-Uncensored-FP8", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "d0xin/Swift-1.5-Qwen3.8-27B-Uncensored-FP8" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "d0xin/Swift-1.5-Qwen3.8-27B-Uncensored-FP8", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use d0xin/Swift-1.5-Qwen3.8-27B-Uncensored-FP8 with Docker Model Runner:
docker model run hf.co/d0xin/Swift-1.5-Qwen3.8-27B-Uncensored-FP8
Related models: all models
Swift-1.5-Qwen3.8-27B-Uncensored-BF16 · Swift-1.5-Qwen3.8-27B-Uncensored-FP8 · Swift-1.5-Qwen3.8-27B-Uncensored-FP8-NInfer · Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE · Swift-1.5-Qwen3.8-Flash-Next-Rank2-Abliteration-Patch · Swift-Qwen3.8-27B-FP8 · Swift-Qwen3.8-27B-Uncensored-BF16 · Swift-Qwen3.8-27B-Uncensored-FP8 · Swift-Qwen3.8-27B-Uncensored-NVFP4-LocalHessian-ActivationHeadroom-NInfer
Swift-1.5-Qwen3.8-27B-Uncensored-FP8
🔓 FP8 quantization of
d0xin/Swift-1.5-Qwen3.8-27B-Uncensored-BF16.The source BF16 checkpoint is an independent uncensored derivative of
ukisai/Swift-1.5-Qwen3.8-27b, produced using rank-1 directional residual-stream ablation.
This release targets substantially lower model memory while retaining the Swift 1.5 architecture, multimodal support, tool calling, MTP weights and long-context configuration.
Initial capability and runtime evaluations are now available: 80.36% on the fixed 280-question MMLU-Pro subset, 85% IFEval prompt-strict / 90.18% instruction-strict, and 143.93 tok/s in the documented local NInfer agentic benchmark on an NVIDIA RTX PRO 6000 Blackwell 96 GB.
Additional refusal-behavior, dedicated multimodal and long-context evaluations remain separate work and will only be reported when validated specifically for this Swift 1.5 FP8 release.
Model lineage
Qwen/Qwen3.8-27B
↓
ukisai/Swift-1.5-Qwen3.8-27b
↓
rank-1 directional residual-stream ablation
↓
d0xin/Swift-1.5-Qwen3.8-27B-Uncensored-BF16
↓
FP8 quantization
↓
d0xin/Swift-1.5-Qwen3.8-27B-Uncensored-FP8
Model summary
- FP8 checkpoint
- approximately 29 GB
- 18 safetensors shards
- configured context: 262,144 tokens
- source ablation layer: 38
- source ablation rank: 1
- modified residual writers: 131
- vision tower unchanged by the ablation
- MTP residual writers included
- multimodal architecture retained
- tool-calling architecture retained
FP8 conversion
The FP8 checkpoint was produced from the validated uncensored BF16 derivative.
Current artifact audit:
- 18 safetensors shards
- 1,606 tensors
- 407
float8_e4m3fntensors - 1,199 BF16 tensors
- 407 scale tensors
- safetensors integrity: PASS
The mixed dtypes are expected: tensors not selected for FP8 quantization remain in BF16 and quantized weights retain their associated scale tensors.
Ablation provenance
The BF16 source was created using:
- source:
ukisai/Swift-1.5-Qwen3.8-27b - ablation layer: 38
- rank: 1
- modified residual writers: 131
- vision tower: unchanged
- MTP residual writers: included
Transformation metadata is included in ABLITERATION.json.
Evaluation status
Initial capability and runtime evaluations are now available for this release.
The results below apply specifically to the Swift 1.5 FP8 derivative in this repository. Results from the earlier Swift-Qwen3.8 release or from the separate Flash-Next derivative are not mixed into this table.
Capability evaluation
| Evaluation | Result |
|---|---|
| MMLU-Pro | 225 / 280 — 80.36% |
| IFEval — prompt strict | 85% |
| IFEval — prompt loose | 86% |
| IFEval — instruction strict | 90.18% |
| IFEval — instruction loose | 91.41% |
The MMLU-Pro evaluation used a fixed 280-question subset consisting of 14 categories × 20 questions.
These results should be treated as reference measurements rather than comprehensive estimates of general model quality.
Agentic / reasoning runtime benchmark
A local xhigh reasoning run was also measured on the NInfer conversion derived from this FP8 checkpoint.
Test system
- NVIDIA RTX PRO 6000 Blackwell 96 GB
- NInfer runtime
- concurrency: 1
- FP8 KV cache
- context / KV capacity: 65,536 tokens
- prefill chunk: 1,024
- DFlash2 speculative decoding
- 5 draft tokens
- LM-head draft enabled
- thinking preserved
Measured run
| Metric | Result |
|---|---|
| Generation throughput | 143.93 tok/s |
| Prefill throughput | 1,625.62 tok/s |
| Completion tokens | 21,764 |
| Reasoning tokens | 15,177 |
| Final-answer tokens | ~6,587 |
| DFlash2 acceptance | 39.9% |
This is a runtime/workload-specific performance result, not a hardware-independent property of the checkpoint. Throughput will vary with inference engine, context length, concurrency, speculative-decoding configuration and hardware.
The corresponding NInfer release is:
d0xin/Swift-1.5-Qwen3.8-27B-Uncensored-FP8-NInfer
Validation coverage
| Area | Status |
|---|---|
| Checkpoint structure | PASS |
| Safetensors integrity | PASS |
| MMLU-Pro | Completed |
| IFEval | Completed |
| Agentic / reasoning performance | Completed |
| Fixed 100-prompt refusal evaluation | Not yet reported for Swift 1.5 FP8 |
| Dedicated multimodal capability evaluation | Not yet reported |
| Dedicated long-context retrieval evaluation | Not yet reported |
An early MATH-500 run is intentionally not reported because the evaluation adapter / answer-format handling for that run was invalid. Its resulting score must not be interpreted as a model capability result.
Further evaluations can be added to this model card without changing the released checkpoint.
Recommended generation settings
A good starting point based on the upstream Swift 1.5 configuration:
| Parameter | Value |
|---|---|
reasoning_effort |
xhigh |
temperature |
1.0 |
top_p |
0.95 |
top_k |
20 |
min_p |
0.0 |
presence_penalty |
0.0 |
repetition_penalty |
1.0 |
SGLang
python -m sglang.launch_server \
--model-path d0xin/Swift-1.5-Qwen3.8-27B-Uncensored-FP8 \
--served-model-name Swift-1.5-Qwen3.8-27B-Uncensored-FP8 \
--trust-remote-code \
--context-length 262144 \
--reasoning-parser qwen3 \
--tool-call-parser qwen3_coder \
--port 30000
Adjust context length and memory parameters for the available hardware.
Tool calling
Use:
--tool-call-parser qwen3_coder
Multimodal support
The Qwen3.8 multimodal architecture and vision tower are retained.
The directional ablation used to create the BF16 source did not modify the vision tower.
Safety and responsible use
This is an uncensored / refusal-reduced derivative model.
The model has been intentionally modified to reduce refusal behavior. It may therefore generate content that the upstream model would normally refuse, restrict or handle more cautiously.
Outputs may be inaccurate, offensive, unsafe, unlawful or otherwise unsuitable for a particular use case.
Users are responsible for evaluating model outputs and ensuring compliance with applicable laws, licenses, regulations and platform requirements.
Upstream models
FP8 source:
d0xin/Swift-1.5-Qwen3.8-27B-Uncensored-BF16
Original Swift 1.5 model:
ukisai/Swift-1.5-Qwen3.8-27b
Base model:
Qwen/Qwen3.8-27B
License
These weights are distributed under the Swift Open License v1.0.
The underlying Qwen3.8 components remain subject to their applicable upstream license.
See the included LICENSE, LICENSE-APACHE-2.0, and NOTICE files for the
applicable terms and attribution requirements.
Attribution
- Base model:
Qwen/Qwen3.8-27B - Swift post-training: UkisAI
- Swift 1.5:
ukisai/Swift-1.5-Qwen3.8-27b - Directional ablation:
d0xin - FP8 quantization and release packaging:
d0xin
Citation
@misc{swift-1.5-qwen3.8-27b,
title = {Swift 1.5 Qwen3.8-27B},
author = {UkisAI},
year = {2026},
url = {https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b}
}
- Downloads last month
- 498