Qwen-Image-2.1 MLX-Serve 4-bit

4-bit pack of Qwen/Qwen-Image-2.1 for mlx-serve: 11.2 GB, for 16 GB Macs.

sample

What is in it

The checkpoint's own diffusers layout and key names, with the DiT block linears and the text-encoder layer linears affine-quantized to 4-bit (group 64). Kept dense: the VAE (f32), embed_tokens, norms, and the DiT's small or shared linears. Kept: the Qwen3-VL vision tower, for instruction editing. Dropped: lm_head and the VAE's per-frame time_convs. Built by tests/convert_qwen_image21_weights.py --preset 16gb.

Measured (M1 Pro, 32 GB)

Pack Size Steps Wall clock incl. load Peak memory
8-bit 1024x1024 40 985 s (~23 s/step) 12.95 GB
4-bit 1024x1024 3 87 s 9.55 GB
4-bit 512x512 20 118 s -

On a Mac the full set would crowd, mlx-serve loads the text encoder per request and frees it before the denoise, so the resident set is the DiT and VAE.

Run it

brew tap ddalcu/mlx-serve https://github.com/ddalcu/mlx-serve
brew install mlx-serve
mlx-serve pull ddalcu/Qwen-Image-2.1-MLX-Serve-4bit
mlx-serve serve
curl localhost:11234/v1/images/generations -H 'Content-Type: application/json' \
  -d '{"model":"ddalcu/Qwen-Image-2.1-MLX-Serve-4bit","prompt":"a red fox in fresh snow","size":"1024x1024"}'

40 steps when steps is omitted. guidance_scale above 1 with a negative_prompt runs real CFG (two forwards per step). image + strength does image-to-image. "mode":"edit" with an image (plus up to 9 ref_images) edits it from the prompt.

Apache-2.0, same as the base model.

Downloads last month
3,255
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for ddalcu/Qwen-Image-2.1-MLX-Serve-4bit

Quantized
(111)
this model

Space using ddalcu/Qwen-Image-2.1-MLX-Serve-4bit 1

Collection including ddalcu/Qwen-Image-2.1-MLX-Serve-4bit