Qwen-Image 2.1 on one RTX 5090: 7x faster, near-lossless
1.85 s per 1024x1024 image on a single RTX 5090 with sglang: 7.3x faster than sglang in BF16 and 4x faster than sglang's own FP8 + SageAttention2 flags, at the image quality of the DPCache paper's accepted settings.
- Code: fractalyze/sglang @
qi21/showcase - This repo: calibrated DPCache schedules, results and samples. No model weights: use Qwen/Qwen-Image-2.1 (its own license applies).
What's inside
- FP8 kernels for RTX 5090 (SM120): block-scaled FP8 GEMMs and fused quantize/activation kernels on top of sglang's FP8 + SageAttention2 path. Same results as sglang's FP8 path.
- DPCache step caching: only 14 of the 40 steps run the full transformer; the rest are extrapolated. The schedule is calibrated for this FP8 stack.
- Faster encode/decode: FP8 VAE decoder and a CUDA-graphed text encoder.
Results
1024x1024, 40 steps, batch 1, RTX 5090, warm latency.
| Configuration | Latency | Speedup |
|---|---|---|
| sglang, BF16 | 13.47 s | 1.0x |
| sglang, FP8 + SageAttention2 | 7.33 s | 1.8x |
| This showcase | 1.85 s | 7.3x |
Quality on 200 held-out DrawBench prompts, compared with sglang BF16 (same prompt and seed):
| Configuration | LPIPS (mean) | PSNR | SSIM | ImageReward change |
|---|---|---|---|---|
| sglang, FP8 + SageAttention2 | 0.087 | 26.9 | 0.913 | +0.014 |
| This showcase | 0.112 | 25.4 | 0.888 | +0.015 |
This is within the DPCache paper's accepted level (mean LPIPS 0.18). On a few prompts the FP8 model draws a different but equally clean composition; sglang's own FP8 flags do the same.
More side-by-side samples (BF16 | showcase) are in samples/.
Quick start
Requires an RTX 5090, CUDA 13 and Python 3.12.
git clone --branch qi21/showcase https://github.com/fractalyze/sglang.git && cd sglang
uv venv -p 3.12 .venv && source .venv/bin/activate
SGLANG_BUILD_RUST_EXTS=none uv pip install -e "python[diffusion]"
uv pip install nvidia-cudnn-frontend==1.29.0
# SageAttention2 for SM120
git clone https://github.com/thu-ml/SageAttention.git ../SageAttention && (cd ../SageAttention && git checkout d1a57a5 && \
TORCH_CUDA_ARCH_LIST=12.0 uv pip install --python ../sglang/.venv/bin/python --no-build-isolation .)
Download the checkpoint and the schedules, then serve:
MODEL=$(hf download Qwen/Qwen-Image-2.1 --revision 790c92633540aa0cb11d9abf19eb46d861714758 --quiet)
hf download Fractalyze/qwen-image21-rtx5090-showcase --include "schedules/*" --local-dir ./showcase
SGLANG_ENABLE_QWEN3VL_TEXT_CUDA_GRAPH=1 SGLANG_ENABLE_QWEN_IMAGE21_VAE_FP8_CONV=1 \
sglang serve --model-path "$MODEL" --port 30000 \
--quantization fp8 \
--attention-backend torch_sdpa --component-attention-backends transformer=sage_attn \
--dit-layerwise-offload false \
--dpcache-schedule-dir ./showcase/schedules --dpcache-default-budget 14
Generate an image:
curl -s http://localhost:30000/v1/images/generations -H "Content-Type: application/json" \
-d '{"prompt": "A red colored car.", "size": "1024x1024", "seed": 42, "response_format": "b64_json", "output_format": "png"}' \
| python -c "import sys, json, base64; open('out.png', 'wb').write(base64.b64decode(json.load(sys.stdin)['data'][0]['b64_json']))"
The first request after startup is slower (one-time warm-up).
Limitations
- The kernels target the RTX 5090 (SM120).
- The schedules are calibrated for this checkpoint revision at 1024x1024, 40 steps; other settings need recalibration.
Model tree for Fractalyze/qwen-image21-rtx5090-showcase
Base model
Qwen/Qwen-Image-2.1