MTP speculative decoding fails with `ValueError: There is no module or parameter named 'draft_lm_head'`
Hi,
Thank you for releasing this model. I'm trying to serve it with vLLM using the recommended MTP configuration, but the engine fails to load the draft model. I've followed the MTP setup as described in the vLLM documentation and the model card, but I consistently hit the same error.
Environment:
- vLLM versions tested:
0.30.1rc1.dev605+g84593955f0.30.1rc1.dev619+g5f30fc703.precompiled(nightly)0.30.0(stable release)
- GPU: 2× RTX 3090 (NVLink, P2P OK)
- Model:
numsu/AREX-2-27B-INT8-W8A16-MTP
Command:
vllm serve numsu/AREX-2-27B-INT8-W8A16-MTP \
--tensor-parallel-size 2 \
--max-model-len 262144 \
--kv-cache-dtype fp8_e4m3 \
--speculative-config '{"method":"mtp","num_speculative_tokens":3}' \
# ... other flags
Error:
The worker crashes during draft model loading with:
ValueError: There is no module or parameter named 'draft_lm_head' in Qwen3_5MultiTokenPredictor.
The available parameters belonging to (Qwen3_5MultiTokenPredictor) are:
{'layers.0.self_attn.attn.v_scale', 'layers.0.self_attn.attn.k_scale', ...}
The full traceback points to vllm/model_executor/models/qwen3_5_mtp.py in load_weights, where remap_weight_names maps the checkpoint's mtp.lm_head.* tensors to model.lm_head.weight, but the Qwen3_5MultiTokenPredictor module does not have a lm_head attribute.
The same error reproduces on all tested versions, including the stable 0.30.0 release and the latest nightly build 0.30.1rc1.dev619+g5f30fc703.precompiled. This suggests it's a general vLLM limitation in handling a separate MTP lm_head, not specific to a single development build.
It appears that the MTP head in this checkpoint is stored as a separate lm_head (i.e., mtp.lm_head.*), but vLLM's current MTP loading path expects the draft head to share the target model's lm_head and does not handle a distinct draft lm_head. I've found a related upstream PR (vllm-project/vllm#52013) that addresses this exact mismatch, but it is not yet in any stable release.
Question:
Is there a known workaround to run this model with MTP on the current vLLM releases? For example:
- Is there a specific vLLM commit or build that handles the separate
draft_lm_headcorrectly? - Should the
--speculative-configuse a different method (e.g.,qwen3_next_mtp) for this architecture? - Or is the only working path currently to serve the model without speculative decoding?
I've also tried method: "dflash" with a DFlash2 drafter (incoai/Qwen3.8-27B-DFlash2), and while it loads and runs, I'd prefer to use the native MTP head if possible.
Any guidance would be greatly appreciated. Thank you!
Thanks for investigating this—and you’re right that the bare command in my model card doesn’t work with the stock vLLM builds you tested. Sorry; the card left out an important runtime requirement.
This checkpoint’s draft head is a 40,960-token shortlist, stored as mtp.draft_lm_head.weight with a separate vocabulary-ID map. Stock vLLM’s Qwen3.5 MTP loader doesn’t create that draft_lm_head module, which is why loading fails. My serving setup uses a custom vLLM patch that adds support for this head: HyperQwen’s Qwen3.5 MTP draft-vocabulary patch.
The method: "mtp" setting is correct; switching to qwen3_next_mtp won’t fix the weight-loading issue. PR #52013 addresses a different layout—mtp.lm_head.*—so it won’t by itself support this checkpoint’s shortlist head.
For this exact checkpoint, use a vLLM build with the shortlist-head patch. Without it, the fallback is to serve without speculative decoding.