Instructions to use ethantodd4l/Swift-1.5-Qwen3.8-27b-FP8 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ethantodd4l/Swift-1.5-Qwen3.8-27b-FP8 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="ethantodd4l/Swift-1.5-Qwen3.8-27b-FP8") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("ethantodd4l/Swift-1.5-Qwen3.8-27b-FP8") model = AutoModelForMultimodalLM.from_pretrained("ethantodd4l/Swift-1.5-Qwen3.8-27b-FP8", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use ethantodd4l/Swift-1.5-Qwen3.8-27b-FP8 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ethantodd4l/Swift-1.5-Qwen3.8-27b-FP8" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ethantodd4l/Swift-1.5-Qwen3.8-27b-FP8", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/ethantodd4l/Swift-1.5-Qwen3.8-27b-FP8
- SGLang
How to use ethantodd4l/Swift-1.5-Qwen3.8-27b-FP8 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "ethantodd4l/Swift-1.5-Qwen3.8-27b-FP8" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ethantodd4l/Swift-1.5-Qwen3.8-27b-FP8", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "ethantodd4l/Swift-1.5-Qwen3.8-27b-FP8" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ethantodd4l/Swift-1.5-Qwen3.8-27b-FP8", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use ethantodd4l/Swift-1.5-Qwen3.8-27b-FP8 with Docker Model Runner:
docker model run hf.co/ethantodd4l/Swift-1.5-Qwen3.8-27b-FP8
Swift-1.5-Qwen3.8-27b-FP8
ukisai/Swift-1.5-Qwen3.8-27b
quantized to FP8 with the same recipe and checkpoint format as the official
Qwen/Qwen3.8-27B-FP8: weights in
FP8 E4M3 with 128x128 block scales, activations dynamically quantized
(activation_scheme: dynamic), serialized in the compressed-tensors/
quant_method: fp8 layout that vLLM loads directly.
Quantization
| property | value |
|---|---|
| weight format | FP8 E4M3, per-128x128-block scales (weight_scale_inv, BF16) |
| activation format | FP8 E4M3, dynamic (no stored input scales) |
| quantized linears | 407 (400 language-decoder + 7 MTP) |
| unquantized | vision tower (model.visual.*), lm_head, norms, embeddings, GDN conv/state parameters |
modules_to_not_convert is a superset of the reference checkpoint's list (1111
vs 882 entries): every module that is BF16 in this checkpoint is listed, and the
quantized tensor set is exactly the reference's.
Verification
- Tensor inventory matched against
Qwen/Qwen3.8-27B-FP8: identical 1606 keys (407F8_E4M3+ 1199BF16), identical quantized module set, no name or shape differences. - Served with vLLM Radiance (vLLM 0.28.0, gfx1201, tensor-parallel 2, R4D attention, prefix caching) and produced correct generations.
Evaluation
Measured with the gsm8k bench of
check-source, the successor to
gsm8k-eval. It reproduces the lm-evaluation-harness gsm8k task exactly:
identical prompt format, per-document 5-shot sampling (seed 1234), filters and
generation settings for every checkpoint, full 1319-item test split.
check-source also carries the chat-template GSM8K protocol (one user turn,
reasoning_effort levels) plus AIME and MATH-500 benches.
Raw result files for these runs:
ethantodd4l/check-source-results.
| checkpoint | thinking strict | thinking flexible | non-thinking strict | non-thinking flexible |
|---|---|---|---|---|
| this model (W8A8 FP8, fp8 KV) | 85.75% | 85.67% ±1.89 | 85.97% | 88.70% ±1.71 |
| Swift-1.5-Qwen3.8-27b bf16 (base) | pending | pending | pending | pending |
Chat protocol (check-source gsm8k-chat)
The raw bench fixes the encoder prompt and the two AMD modes; the chat bench puts the same lm-eval 5-shot text in a single user turn and lets the checkpoint's template choose the thinking level, with the reasoning-level sampling defaults (T=1.0, top_p=0.95, top_k=20). Same checkpoint family, only the serving configuration changes (200 items, seed 1234, target-only, no speculation):
| weights | activations | KV cache | none | medium | xhigh | mean tokens none/med/xhigh |
|---|---|---|---|---|---|---|
| bf16 (base) | bf16 | fp16 | 97.5% | 98.5% | 98.5% | 178 / 328 / 427 |
| FP8 W8A8 (this model) | fp8 dynamic per-token | fp16 | 98.5% | 98.0% | 99.5% | 162 / 328 / 452 |
| FP8 W8A8 (this model) | fp8 dynamic per-token | fp8 | 97.0% | 98.0% | 98.0% | 168 / 305 / 481 |
| MXFP4 RTN | fp8 WMMA (W4A8) | fp16 | 96.5% | 98.0% | 97.0% | 148 / 317 / 500 |
| MXFP4 RTN | fp8 WMMA (W4A8) | fp8 | 96.5% | 99.0% | 96.5% | 162 / 308 / 497 |
Scores are flexible-extract; at 200 items the 95% interval is roughly 1.5 points, so the accuracy differences above sit inside noise. The consistent signals are the KV direction (fp16 beats fp8 in every pair) and the token counts. A full 1319-item run plus AIME/MATH-500 is the next step.
Activation and KV-cache precision
Dequantized to bf16 (undoing the 128x128 block scales), this checkpoint is
nearly lossless: next-token logits for "The capital of France is" against the
bf16 base give rel=0.024, corr=0.999 and the same top-1 token (Paris). The
GSM8K gap above therefore comes from the runtime numerics, not the stored
weights:
- FP8 activations (W8A8 dynamic, AITER per-token) are the dominant accuracy cost on this stack. Weight-only (W8A16) would remove it, but this vLLM build has no ROCm W8A16 kernel; it needs a radiance-side kernel. The ~9-point gap above is the full-split raw protocol; on the first 200 items with the chat protocol this same W8A8 configuration is within noise of the bf16 base, so the size of the gap is protocol- and level-dependent.
- fp8 KV cache costs about 1.3 points by itself in a 300-item seeded A/B
(80.0% -> 81.3% flexible). Switching to fp16 KV (
--kv-cache-dtype auto) does not recover the bulk of the gap. - Production guidance: use fp16 KV for accuracy (at ~half the KV capacity), and prefer the MXFP4 RTN sibling for this base model until a W8A16 kernel exists.
Serving
vllm serve ethantodd4l/Swift-1.5-Qwen3.8-27b-FP8 \
--served-model-name swift-1.5-fp8 \
--tensor-parallel-size 2 \
--max-model-len 16384 \
--kv-cache-dtype auto
The FP8 lane is the default of
vLLM Radiance (ROCm/gfx1201,
vLLM 0.28.0); the MXFP4-specific switches (RADIANCE_MXFP4,
RADIANCE_MXFP4_W4A8, RADIANCE_QUARK_BF16_MTP) do not apply to this
checkpoint. The evaluation row above was measured with fp8 KV; auto (fp16)
recovers about 1.3 points in a 300-item probe, though not the bulk of the W8A8
gap.
License
The quantization is a derivative of
ukisai/Swift-1.5-Qwen3.8-27b
and is distributed under the Swift Open License v1.0 (see LICENSE). The
Qwen3.8-27B base model and the tokenizer/config files retained from it remain
under their original Apache License 2.0 terms; see NOTICE.
- Downloads last month
- 53