Instructions to use Iwannapose/minimax_h3_pdmd_4nfe_comfyui with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use Iwannapose/minimax_h3_pdmd_4nfe_comfyui with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline from diffusers.utils import export_to_video # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("MiniMaxAI/MiniMax-H3", dtype=torch.bfloat16, device_map="cuda") pipe.load_lora_weights("Iwannapose/minimax_h3_pdmd_4nfe_comfyui") prompt = "A man with short gray hair plays a red electric guitar." output = pipe(prompt=prompt).frames[0] export_to_video(output, "output.mp4") - Inference
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Draw Things
Download README.md from Iwannapose/minimax_h3_pdmd_4nfe_comfyui: direct link, hf CLI and curl.
- Browser
- Download file 2.62 kB
-
https://huggingface.co/Iwannapose/minimax_h3_pdmd_4nfe_comfyui/resolve/main/README.md
- Command line
-
hf download hf://Iwannapose/minimax_h3_pdmd_4nfe_comfyui/README.md
-
curl -L -o README.md https://huggingface.co/Iwannapose/minimax_h3_pdmd_4nfe_comfyui/resolve/main/README.md
license: apache-2.0
library_name: diffusers
base_model: MiniMaxAI/MiniMax-H3
pipeline_tag: text-to-video
tags:
- video-generation
- distillation
- pdmd
- lora
- comfyui
PDMD 4-NFE LoRA for MiniMax-H3 β ComfyUI format
ComfyUI-format conversion of the PDMD 4-NFE LoRA
(lora_model_0.safetensors) β the 4-step (4 NFE) student distilled from MiniMax-H3-33B with
Projected Distribution Matching Distillation, rank 128, covering attention projections and both
feed-forward layers of all 50 transformer blocks + 2 token-refiner blocks.
Files
| file | description |
|---|---|
minimax_h3_pdmd_4nfe_comfyui_v6.safetensors |
Use this one. Correct ComfyUI conversion (bf16). |
Earlier uploads (base / keys_v2 / v3 / v4 / v5) were broken conversions and have been removed: most of their keys matched nothing in the ComfyUI loader (silent no-op), and v4 missed the SwiGLU gate/value remap below.
What the conversion does
Source keys are Diffusers PEFT names (transformer.transformer_blocks.N.attn.to_q.lora_A.weight β¦);
target keys follow the ComfyUI H3 layout (same scheme as the proven
minimax_h3_fl2v_turbo_8step_v1.0_768p_comfyui_bf16.safetensors):
- q/k/v β fused
qkv_proj:A = [A_q; A_k; A_v](rows),B = block_diag(B_q, B_k, B_v), row order[q; k; v](A[384, 5376], B[21504, 384]). - SwiGLU remap: diffusers
SwiGLUoutputs[value; gate], ComfyUI's H3_swiglu_eagerexpects[gate; up]β the two 14336-row halves of everymlp.fc1.lora_Bare swapped. - alpha entries:
qkv_proj = 384(3 Γ 128),out_proj/fc1/fc2 = 128, so ComfyUI'salpha/rankscale = 1.0 = the PDMD fusion scale (W += (B @ A), alpha/rank = 128/128). to_out.0 β attn.out_proj,ff.net.0.proj β mlp.fc1,ff.net.2 β mlp.fc2;transformer_blocks.N β blocks.N,token_refiner.refiner_blocks.N β token_refiner.blocks.N.
Usage (ComfyUI)
- Node: LoraLoaderModelOnly (model part only β the H3 text encoder is loaded separately).
- Strength: 1.0 (any other value scales the distillation delta, which changes the 4-step behavior β do not use it as a partial-style LoRA).
- Sample at 4 denoising steps (the student's operating point); H3 scheduler config (shift 12 video / 3 audio), no CFG (H3 is guidance-distilled).
- Works on the stock ComfyUI H3 loader; verified that all 208 target keys exist with matching
dimensions in
minimax_h3_ref2va_int8_convrot.safetensors, and every tensor is numerically identical to the original under the transforms above.