Iwannapose's picture
v6: correct ComfyUI conversion (SwiGLU gate/value swap + fused qkv + alpha); remove broken attempts
ae6ab62 verified
|
Raw History Blame Contribute Delete
2.62 kB
---
license: apache-2.0
library_name: diffusers
base_model: MiniMaxAI/MiniMax-H3
pipeline_tag: text-to-video
tags:
- video-generation
- distillation
- pdmd
- lora
- comfyui
---
# PDMD 4-NFE LoRA for MiniMax-H3 β€” ComfyUI format
ComfyUI-format conversion of the [PDMD 4-NFE LoRA](https://huggingface.co/pdmd2026/pdmd_4NFE_lora)
(`lora_model_0.safetensors`) β€” the 4-step (4 NFE) student distilled from MiniMax-H3-33B with
Projected Distribution Matching Distillation, rank 128, covering attention projections and both
feed-forward layers of all 50 transformer blocks + 2 token-refiner blocks.
## Files
| file | description |
|---|---|
| `minimax_h3_pdmd_4nfe_comfyui_v6.safetensors` | **Use this one.** Correct ComfyUI conversion (bf16). |
Earlier uploads (base / keys_v2 / v3 / v4 / v5) were broken conversions and have been removed:
most of their keys matched nothing in the ComfyUI loader (silent no-op), and v4 missed the
SwiGLU gate/value remap below.
## What the conversion does
Source keys are Diffusers PEFT names (`transformer.transformer_blocks.N.attn.to_q.lora_A.weight` …);
target keys follow the ComfyUI H3 layout (same scheme as the proven
`minimax_h3_fl2v_turbo_8step_v1.0_768p_comfyui_bf16.safetensors`):
1. **q/k/v β†’ fused `qkv_proj`**: `A = [A_q; A_k; A_v]` (rows), `B = block_diag(B_q, B_k, B_v)`,
row order `[q; k; v]` (A `[384, 5376]`, B `[21504, 384]`).
2. **SwiGLU remap**: diffusers `SwiGLU` outputs `[value; gate]`, ComfyUI's H3 `_swiglu_eager`
expects `[gate; up]` β€” the two 14336-row halves of every `mlp.fc1.lora_B` are swapped.
3. **alpha entries**: `qkv_proj = 384` (3 Γ— 128), `out_proj/fc1/fc2 = 128`, so ComfyUI's
`alpha/rank` scale = **1.0** = the PDMD fusion scale (`W += (B @ A)`, alpha/rank = 128/128).
4. `to_out.0 β†’ attn.out_proj`, `ff.net.0.proj β†’ mlp.fc1`, `ff.net.2 β†’ mlp.fc2`;
`transformer_blocks.N β†’ blocks.N`, `token_refiner.refiner_blocks.N β†’ token_refiner.blocks.N`.
## Usage (ComfyUI)
- Node: **LoraLoaderModelOnly** (model part only β€” the H3 text encoder is loaded separately).
- **Strength: 1.0** (any other value scales the distillation delta, which changes the 4-step
behavior β€” do not use it as a partial-style LoRA).
- Sample at **4 denoising steps** (the student's operating point); H3 scheduler config
(shift 12 video / 3 audio), no CFG (H3 is guidance-distilled).
- Works on the stock ComfyUI H3 loader; verified that all 208 target keys exist with matching
dimensions in `minimax_h3_ref2va_int8_convrot.safetensors`, and every tensor is numerically
identical to the original under the transforms above.