
Qwen3.8-Flash-Next-Sushi-4bpw
Qwen/Qwen3.8-Flash-Next packed for sushi, a native inference engine for Apple Silicon. Target: a 96 GB Mac. It does not fit on a 64 GB Mac; use Sushi-3bpw there.
| part | how it is stored |
|---|---|
| routed experts (48 layers x 512, plus the MTP layer) | EXL3, made with our own converter |
| dense part (attention, GDN, shared experts, hyper-connections, indexer, embedding, lm_head) | 8-bit affine; the same in both packs |
| n-gram embedding table | bf16 as released, one file ngram_table.bin, read from the SSD and never loaded into GPU memory (Sushi-3bpw ships a 4-bit copy) |
| vision tower (Qwen3-VL ViT) | included: image and video input |
| MTP head | included |
| context | up to 1,048,576 tokens (config.json sets YaRN x4 over the native 262,144) |
| disk | ~159 GiB: 63.7 GiB of weights + 95.4 GiB n-gram table |
sushi only. This pack needs sushi v1.0.0 or later on Apple Silicon with macOS 26.2 or later. It does not load in transformers, vLLM, mlx-lm or exllamav3.
Run it on a 96 GB Mac
# sushi: the release binary (ad-hoc signed; curl does not quarantine it)
curl -L https://github.com/beamivalice/sushi/releases/latest/download/sushi-bin-macos-arm64.tar.gz | tar xz
# this pack
hf download beamster/Qwen3.8-Flash-Next-Sushi-4bpw --local-dir ~/.sushi/models/Qwen3.8-Flash-Next-Sushi-4bpw
# let the GPU use 88000 MB of the 96 GB (until the next reboot)
sudo sysctl iogpu.wired_limit_mb=88000
# serve on 127.0.0.1:12345
./sushi-macos-arm64/sushi serve --model ~/.sushi/models/Qwen3.8-Flash-Next-Sushi-4bpw --mtp-head-kv-quant
- MTP (the model's own draft head) and the 8-bit KV cache are on by default;
--mtp-head-kv-quantstores the MTP head's own KV at 8 bits too. - A browser download of the binary is quarantined by macOS: clear it with
xattr -dr com.apple.quarantine sushi-macos-arm64.
Memory
GPU memory in GiB to serve one prompt that fills the whole context (8-bit KV, MTP on, --mtp-head-kv-quant,
--prefix-cache-mem 1GB). The n-gram table stays on the SSD and is not counted.
| context | Sushi-2bpw | Sushi-2.6bpw | Sushi-3bpw | Sushi-4bpw |
|---|---|---|---|---|
| weights only | 35.0 | 44.0 | 49.3 | 63.7 |
| 128k | 41.7 | 50.7 | 56.1 | 70.4 |
| 256k | 44.6 | 53.6 | 58.9 | 73.3 |
| 512k | 49.7 | 58.6 | 64.0 | 78.4 |
| 1M | 59.8 | 68.8 | 74.2 | 88.5 |
A context fits when its number is below the GPU limit you set with sudo sysctl iogpu.wired_limit_mb. Max context is
the largest one that fits, at 8-bit / 4-bit KV, with 256 MiB spare and capped at 1M:
| Mac | GPU limit | Sushi-2bpw | Sushi-2.6bpw | Sushi-3bpw | Sushi-4bpw |
|---|---|---|---|---|---|
| 48 GB | 43,000 MB (42.0 GiB) | 128k / 192k | — | — | — |
| 64 GB | 59,000 MB (57.6 GiB) | 896k / 1M | 440k / 744k | 184k / 288k | — |
| 96 GB | 88,000 MB (85.9 GiB) | 1M / 1M | 1M / 1M | 1M / 1M | 880k / 1M |
| 128 GB | 120,000 MB (117.2 GiB) | 1M / 1M | 1M / 1M | 1M / 1M | 1M / 1M |
The launch above leaves --prefix-cache-mem at its default, one session at the working context and never under 2 GB,
so it needs at least 1 GiB more than these numbers; pass --prefix-cache-mem 1GB to match them.
Quality


KLD against the bf16 checkpoint: 16 prompts x 512 tokens, scored to the first EOS (7186 positions), KV cache at kv8, every pack run by sushi: Sushi-3bpw and Sushi-4bpw on a later build than the rest, Sushi-2bpw on v1.0.4 with bf16 KV. Lower is better. The n-gram table is not counted: it stays on the SSD.
| pack | weights in GPU memory (GiB) | KLD | top-1 agreement |
|---|---|---|---|
| Sushi-4bpw | 63.68 | 0.0592 | 92.89% |
| oMLX oQ5e | 83.97 | 0.0625 | 92.40% |
| mlx-serve mixed-4-8bit | 70.13 | 0.0818 | 91.39% |
| Sushi-3bpw (4-bit n-gram table) | 49.33 | 0.1036 | 90.31% |
| Sushi-2.6bpw (4-bit n-gram table) | 43.95 | 0.1355 | 89.08% |
| oMLX oQ4e | 69.21 | 0.1370 | 88.87% |
| mlx-serve iQ-MLX 3.3bpw | 50.60 | 0.1987 | 86.28% |
| Vontra 4-bit g32 (TensorFold) | 75.60 | 0.2074 | 85.01% |
| Sushi-2bpw (4-bit n-gram table) | 34.97 | 0.2080 | 85.94% |
At about the same memory, Sushi-4bpw has less than half oQ4e's KLD (0.059 vs 0.137), and it beats oQ5e (0.059 vs 0.063) with 20 GiB less. At about 50 GiB, Sushi-3bpw has about half the KLD of mlx-serve's iQ-MLX 3.3bpw (0.104 vs 0.199). Each pack is measured with the n-gram table it ships: 4-bit g32 in Sushi-2bpw, Sushi-2.6bpw, Sushi-3bpw and mlx-serve's packs, 4- or 5-bit in oMLX's, bf16 in Sushi-4bpw. The oMLX and mlx-serve packs were measured as published; the oMLX n-gram tables were repacked into sushi's file unchanged.
Hard questions

LLMProbe's Hard 50 set (MMLU-Pro, OlympiadBench, LiveBench, NIST Juliet): 45 correct, 5 incorrect, 0 partial, 90.0%. llmprobe 0.6.12, xhigh effort, 12k reasoning / 16k output tokens, 8-bit KV and MTP.
Choosing the n-gram table
Sushi-3bpw and Sushi-4bpw ship different copies of the same n-gram table, and either copy works with either pack. Use the 4-bit table to save 66 GiB of disk, or the bf16 table for a small quality gain. The table is read from the SSD, so GPU memory is the same either way.
| table | size on disk | ships with | KLD with Sushi-3bpw | KLD with Sushi-4bpw |
|---|---|---|---|---|
| 4-bit, group size 32 | 29.8 GiB | Sushi-3bpw | 0.1036 | 0.0654 |
| bf16 | 95.4 GiB | Sushi-4bpw | 0.1006 | 0.0592 |
To swap, download the other pack's ngram_table.bin into this pack's folder, replacing the current one:
hf download beamster/Qwen3.8-Flash-Next-Sushi-3bpw ngram_table.bin --local-dir <this pack's folder>
Then set ngram_table in config.json to match the file:
"ngram_table": {"file": "ngram_table.bin", "bits": 4, "group_size": 32} for the 4-bit table.
Credits and license
- Qwen team: the base model, Qwen/Qwen3.8-Flash-Next. This
pack is a derivative under the Qwen Community License 1.0 (
LICENSE), including its conditions on large commercial services. - turboderp: the EXL3 format.
- ddalcu: mlx-serve. sushi is a fork of it.
- David Tai: his work on --mtp-typical
- Downloads last month
- 392
Model tree for beamster/Qwen3.8-Flash-Next-Sushi-4bpw
Base model
Qwen/Qwen3.8-Flash-Next