Qwen-Image 2.1 on one RTX 5090: 7x faster, near-lossless

1.85 s per 1024x1024 image on a single RTX 5090 with sglang: 7.3x faster than sglang in BF16 and 4x faster than sglang's own FP8 + SageAttention2 flags, at the image quality of the DPCache paper's accepted settings.

What's inside

  • FP8 kernels for RTX 5090 (SM120): block-scaled FP8 GEMMs and fused quantize/activation kernels on top of sglang's FP8 + SageAttention2 path. Same results as sglang's FP8 path.
  • DPCache step caching: only 14 of the 40 steps run the full transformer; the rest are extrapolated. The schedule is calibrated for this FP8 stack.
  • Faster encode/decode: FP8 VAE decoder and a CUDA-graphed text encoder.

Results

1024x1024, 40 steps, batch 1, RTX 5090, warm latency.

Configuration Latency Speedup
sglang, BF16 13.47 s 1.0x
sglang, FP8 + SageAttention2 7.33 s 1.8x
This showcase 1.85 s 7.3x

Quality on 200 held-out DrawBench prompts, compared with sglang BF16 (same prompt and seed):

Configuration LPIPS (mean) PSNR SSIM ImageReward change
sglang, FP8 + SageAttention2 0.087 26.9 0.913 +0.014
This showcase 0.112 25.4 0.888 +0.015

This is within the DPCache paper's accepted level (mean LPIPS 0.18). On a few prompts the FP8 model draws a different but equally clean composition; sglang's own FP8 flags do the same.

LPIPS distribution

More side-by-side samples (BF16 | showcase) are in samples/.

Quick start

Requires an RTX 5090, CUDA 13 and Python 3.12.

git clone --branch qi21/showcase https://github.com/fractalyze/sglang.git && cd sglang
uv venv -p 3.12 .venv && source .venv/bin/activate
SGLANG_BUILD_RUST_EXTS=none uv pip install -e "python[diffusion]"
uv pip install nvidia-cudnn-frontend==1.29.0
# SageAttention2 for SM120
git clone https://github.com/thu-ml/SageAttention.git ../SageAttention && (cd ../SageAttention && git checkout d1a57a5 && \
  TORCH_CUDA_ARCH_LIST=12.0 uv pip install --python ../sglang/.venv/bin/python --no-build-isolation .)

Download the checkpoint and the schedules, then serve:

MODEL=$(hf download Qwen/Qwen-Image-2.1 --revision 790c92633540aa0cb11d9abf19eb46d861714758 --quiet)
hf download Fractalyze/qwen-image21-rtx5090-showcase --include "schedules/*" --local-dir ./showcase

SGLANG_ENABLE_QWEN3VL_TEXT_CUDA_GRAPH=1 SGLANG_ENABLE_QWEN_IMAGE21_VAE_FP8_CONV=1 \
sglang serve --model-path "$MODEL" --port 30000 \
  --quantization fp8 \
  --attention-backend torch_sdpa --component-attention-backends transformer=sage_attn \
  --dit-layerwise-offload false \
  --dpcache-schedule-dir ./showcase/schedules --dpcache-default-budget 14

Generate an image:

curl -s http://localhost:30000/v1/images/generations -H "Content-Type: application/json" \
  -d '{"prompt": "A red colored car.", "size": "1024x1024", "seed": 42, "response_format": "b64_json", "output_format": "png"}' \
  | python -c "import sys, json, base64; open('out.png', 'wb').write(base64.b64decode(json.load(sys.stdin)['data'][0]['b64_json']))"

The first request after startup is slower (one-time warm-up).

Limitations

  • The kernels target the RTX 5090 (SM120).
  • The schedules are calibrated for this checkpoint revision at 1024x1024, 40 steps; other settings need recalibration.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Fractalyze/qwen-image21-rtx5090-showcase

Finetuned
(53)
this model

Paper for Fractalyze/qwen-image21-rtx5090-showcase