MiMo-V2.6-Flash (MOPD) on 2Γ RTX PRO 6000 Blackwell: FP8 KV cache, W4A8 MoE, custom sm_120 kernels
A serving recipe that runs MiMo-V2.6-Flash-MOPD (recommended) or MiMo-V2.6-Flash-RL (309B total / 15B active, MXFP4 experts, text + image + video + audio in) on two RTX PRO 6000 Blackwell cards (sm_120, 96 GB each, PCIe, no NVLink) with vLLM at tensor-parallel 2, at 2.1Γ the decode speed and 1.6Γ the prefill speed of the stock vLLM image on the same hardware, with 256K context and all four input modes enabled. The weights are Xiaomi's official checkpoint (already MXFP4 experts + FP8 dense) and are not redistributed here; this repository is the runtime quantization and the code around it:
- FP8 (E4M3) KV cache for the model's DiffKV attention layers (stock vLLM only allows bf16 there). Doubles the KV pool.
- W4A8-FP8 MoE: the MXFP4 experts with FP8 activations on Marlin (
VLLM_MARLIN_INPUT_DTYPE=fp8). +15% prefill, same quality. - Three attention fixes to vLLM's Triton DiffKV kernel (split-KV for the speculative-decode verify step, wide prefill tiles, fp8 K/V), and a purpose-built CUDA prefill attention kernel for the 9 global layers (1.34Γ the Triton kernel).
- A custom small-batch MoE decode kernel on b12x's FP4 layout (beats Marlin at 8β32 routed tokens; kept as an opt-in hybrid backend), the DFlash drafter at 3 draft tokens, and vLLM's CPU KV tier for multi-agent fan-out.
Everything needed to reproduce it is in recipe/ (see recipe/README.md): image build, launcher, vLLM patches, kernels,
tests and the benchmark scripts. Raw benchmark logs are in results/.
Benchmarks
Two RTX PRO 6000 Blackwell Max-Q at their 300 W cap, PCIe, TP2, one 20-agent workload shape: β46K-token prompts at a 61:1
prefill:decode ratio. A β2.5 GB/GPU side process stayed resident during every run. recipe/bench/bench-sbs.py sends
non-streaming requests; decode speed is the difference between a 400-token and a 16-token reply on a warm prefix cache.
bench-fanout.py runs 20 agents Γ 3 turns (β40K tokens each) through 8 slots. GSM8K is the first 200 test problems, greedy.
Base = the stock vllm/vllm-openai:mimo-v26-x86_64-cu130 image with the vendor recipe's settings (Marlin W4A16,
DFlash with 7 draft tokens) as far as they fit on two cards (text-only, bf16 KV, 262K context). Each row adds one change.
| Configuration | 46K prefill, 1 stream (tok/s) | TTFT 46K | 46K decode, 1 stream | 46K, 4 streams: prefill / decode | 100K, 4 streams: prefill / decode | 20-agent fan-out | GSM8K-200 |
|---|---|---|---|---|---|---|---|
| Base (stock image) | 6,230 | 7.4 s | 87β94 | 12,588 / 263 | β | 346 s | 98.0% |
| + split-KV for the verify step | 6,130 | 7.5 s | 156β164 | 12,340 / 335 | β | 344 s | 98.5% |
| + 3 draft tokens instead of 7 | 6,160 | 7.4 s | 168β175 | 12,320 / 386 | β | β | β |
| + wide prefill tiles, CPU KV tier | 8,460 | 5.4 s | 158β175 | 16,873 / 385 | β | 97.9 s | 98.5% |
| + all input modes, FP8 KV cache | 8,210 | 5.6 s | 180β188 | 16,413 / 397 | β | 92.0 s | 98.5% |
| + one program per whole verify | 8,337 | 5.5 s | 187β191 | 16,608 / 412 | 13,334 / 348 | 92.3 s | 99.0% |
| + W4A8-FP8 MoE (Marlin) | 9,532 | 4.8 s | 184β199 | 18,964 / 411 | 14,757 / 367 | 80.7 s | 98.5% |
| + custom prefill attention kernel (this repo's default) | 10,177 | 4.5 s | 189β193 | 20,514 / 416 | 16,840 / 366 | 76.8 s | 98.0% |
At the default config the KV pool is 479K tokens (fp8, all encoders loaded); the 29K-token and 336K-token needle tests, tool calls, and image/video/audio probes pass. Decode is flat from 2K to 46K context and 131 tok/s at 180K.
Things that were measured and did not help on this hardware (logs in results/): b12x MXFP4 MoE alone (+15% prefill,
β45% four-stream decode), MTP instead of DFlash, 8192-token prefill chunks (KV pool 467K β 293K), FlashInfer/custom
all-reduce (32 MB reductions are PCIe-bound; NCCL is at the link's floor), FlashInfer's SM120 192Γ128 fused-attention
kernels (unreachable through the 0.6.18/0.7.0 wrappers), FP8 QKα΅ in the attention kernel (β8% time, 7Γ the error), and
2 CTAs/SM for it (spills).
Update 2026-10-01: three correctness fixes
- Tool-call string arguments are now passed through verbatim. The day-0 image maps
--tool-call-parser mimoto vLLM's Qwen3 parser, which strips one leading and one trailing newline from every string argument. MiMo writes<parameter=β¦>value</parameter>with no newline wrapper, so a file-write call whose content ends in a newline lost it.recipe/docker/build.shnow applies vLLM #58019 (merged upstream 2026-09-29), which adds a MiMo parser. In a live check, 6 of 6 calls kept the trailing newline, against 0 of 6 before. - No NaN for query rows without visible keys. In the split-KV attention reducer, a row whose segments are all empty (for example
the padded rows of a CUDA-graph batch) computed
exp(-inf - -inf)and produced NaN. A one-line clamp, which the upstream port of this patch (vLLM #59085) also has, makes those rows 0.recipe/patches/test_diffkv_emptyrows.pyshows 16 of 16 empty rows NaN before and 0 after; normal rows are unchanged. - Images with large uniform regions are read correctly. The day-0 image adds the vision encoder's attention sink as a bias on
key 0 (
sinks_bias_key0=True), which misreads uniform image regions; a solid green square came back as "a close-up of an eye".recipe/patches/mimo_v2_omni.pyapplies vLLM #58235 (still open upstream), which treats the sink as an extra softmax-denominator logit, plus the omni wrapper'spacked_modules_mapping. Solid green, white and blue squares are now described correctly, and the image, video and audio probes pass.VIT_SINK_FIX=0restores the stock file.
Same-session A/B of the first two fixes, every other setting unchanged (the vision fix was on in both arms;
results/fixes-1001-ab.txt):
| before | after | |
|---|---|---|
| 46K prefill / TTFT | 10.2K / 4.5 s | 10.2K / 4.5 s |
| 46K decode, 1 stream | 179β181 | 186β190 |
| 46K, 4 streams: prefill / decode | 20.4K / 426 | 20.5K / 427 |
| Agent turn (34K cached + 6K new), 1 stream: TTFT / decode | 0.70 s / 205 | 0.71 s / 207 |
| 20-agent fan-out | 77.4 s | 77.0 s |
| GSM8K-200 / HumanEval | 98.0% / 93.3% | 97.0% / 95.1% |
The differences are within run-to-run noise. A torch-profiler run of the fixed build (results/profiles/eco1001-cand.prof.txt)
matches the earlier profile: 13.4 ms per decode step at 46K, GPU busy 94%.
Update 2026-09-28: MiMo-V2.6-Flash-MOPD
Xiaomi's MOPD checkpoint (released 2026-09-27; aimed at tool-call repetition in agent harnesses) has the same config and DFlash drafter as
Flash-RL, so the recipe runs it unchanged: set MODEL to the MOPD download. Same session, same settings (results/rl-0928*.txt,
results/mopd-0928*.txt):
| Flash-RL | MOPD | |
|---|---|---|
| 46K prefill / TTFT | 10.0β10.2K / 4.5β4.6 s | 10.1K / 4.5β4.6 s |
| 46K decode, 1 stream | 184β185 | 180β182 |
| 46K, 4 streams: prefill / decode | 20.6K / 419 | 20.2K / 430 |
| 20-agent fan-out | 75.5 s | 77.1 s |
| GSM8K-200 | 97.5% | 98.5% |
| needle 4.6K/29K, tool calls, image/video/audio | pass | pass |
The benchmark table above was measured on Flash-RL; MOPD is within noise of it on every row we re-ran. This repository's name predates the MOPD release; it serves both checkpoints.
Update 2026-09-24
- Exact QKV weights. The fused-FP8-QKV loader in vLLM re-quantizes each rank's rows below TP4. Measured on this checkpoint at TP2,
that changed 18% of Q, ~51% of K and 73β81% of V weights in the global attention layers (1.3β2.5% relative L2). The recipe now carries
a port of local-inference-lab/vllm #874 that keeps every weight on its checkpoint scale (0 weights changed; audit in
recipe/patches/test_mimo_exact_qkv_audit.py).VLLM_MIMO_EXACT_QKV=0restores the old loader. - Sampling defaults.
serve.shnow uses--generation-config auto: requests without sampling parameters get the checkpoint's temperature 1.0 / top_p 0.95 instead of vLLM's near-greedy defaults, which users report make MiMo repeat the same tool call in agent harnesses. Clients that send no sampling parameters see lower speculative acceptance at temperature 1.0 (mean accept length 2.93 vs 3.42 across our suite); send your owntemperatureor setGEN_CONFIG=vllmif you prefer the old behaviour. - Same-day before/after (
results/fix0924-old.txt,results/fix0924-new.txt): 46K prefill 10.1K / 10.0K tok/s, TTFT 4.6 s both; 46K decode 164 / 185β194 tok/s (noise); four streams at 46K 20.0K / 420; 20-agent fan-out 77.0 s; GSM8K-200 97.5% both; KV pool 479K β 467K tokens (the padded QKV layout). Needle, tool-call and image/video/audio probes pass. A new multi-turn tool-loop probe (recipe/bench/test-mimo-agentic.py) showed no reasoning leaks and no repeated tool calls in 8 runs with either configuration.
What the kernels do
DiffKV attention on SM120. MiMo's K and V head dims differ (192/128), so vLLM routes every target layer to its generic
Triton "DiffKV" kernel (the FlashAttention path needs FA3/FA4). Profiling showed that kernel was 59% of decode time at 46K:
it disables its split-KV mode whenever a request has more than one query token, and every speculative-decode verify has
several, so the 9 full-context layers ran on β18 CTAs of a 188-SM GPU. recipe/patches fixes that (5β7Γ per call), adds
wide prefill tiles (3Γ), and ports the E4M3 K/V path from the local-inference-lab fork with SM120 shared-memory fixes.
recipe/kernels/attention/prefill_attn.cu replaces the Triton kernel for prefill rows on the global layers: FA2-style,
one CTA per 8 tokens Γ 16 GQA heads, ldmatrix + mma.m16n8k16, fp8 K/V converted in the load path. 10.5 ms vs 14.1 ms per
4096-token chunk at 26K context; Nsight shows it compute-bound (tensor pipe 55%, DRAM 0.4%).
recipe/kernels/moe/decode_moe.cu streams b12x's prepared MXFP4 layout at 92% of DRAM bandwidth for decode-sized batches
and feeds decoded FP4 straight into tensor-core fragments for 8β32 tokens. It exists so that b12x (the fastest prefill MoE
kernel on sm_120) and a Marlin-class decode kernel can share one weight copy; the hybrid backend is wired but not yet the
default.
Credits and license
Model: Xiaomi (MIT). Serving stack: vLLM, FlashInfer, Marlin, b12x.
The fp8-KV, cache-config and mixed-batch changes were ported from pull requests in
local-inference-lab/vllm. The code in this repository is MIT
(LICENSE); see THIRD_PARTY_NOTICES.md. Built by Diffbot for a two-card agentic-coding workstation.
Model tree for diffbot/MiMo-V2.6-Flash-MOPD-FP8KV-W4A8-2x-RTX-PRO-6000
Base model
XiaomiMiMo/MiMo-V2.6-Flash-MOPD