Text-to-Video
Diffusers
Safetensors
MiniMaxH3ModularPipeline
video-generation
audio
fastvideo
bf16
pruned
Instructions to use FastVideo/FastVideo-FastH3-Trim-8-Step with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use FastVideo/FastVideo-FastH3-Trim-8-Step with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("FastVideo/FastVideo-FastH3-Trim-8-Step", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
Release model card for FastH3 Trim
Browse files
README.md
CHANGED
|
@@ -1,12 +1,23 @@
|
|
| 1 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 2 |
|
| 3 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
| 4 |
|
| 5 |
-
|
| 6 |
-
|
| 7 |
-
|
| 8 |
-
-
|
| 9 |
-
- Sampling contract: video/audio scheduler shift 10/3, guidance 1.0, VSA_sparsity 0.8, VSA_tile_size 64.
|
| 10 |
-
- `text_encoder` is not included; it is identical to the one in `FastVideo/FastVideo-FastH3-8-Step-V2`.
|
| 11 |
|
| 12 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: other
|
| 3 |
+
base_model: FastVideo/FastVideo-FastH3-8-Step-V2
|
| 4 |
+
tags: [video-generation, text-to-video, audio, fastvideo, bf16, pruned]
|
| 5 |
+
---
|
| 6 |
+
# FastH3 Trim, 8-step, BF16 source weights
|
| 7 |
|
| 8 |
+
FastH3 Trim is an experimental, smaller version of [FastH3 V2](https://huggingface.co/FastVideo/FastVideo-FastH3-8-Step-V2):
|
| 9 |
+
eight-step text-to-video with synchronized audio from 42 of the 50 MiniMax H3 transformer blocks. We removed the eight
|
| 10 |
+
blocks whose removal changed video and audio predictions the least, replaced each block's AdaLN timestep projection with
|
| 11 |
+
a shared rank-16 basis, and trained the result with eight-step DMD2 (checkpoint 300). Removing blocks makes the model
|
| 12 |
+
faster and smaller but costs some quality; use FastH3 V2 when quality matters most.
|
| 13 |
|
| 14 |
+
- **Sampling:** 8 DMD steps (999, 874, 749, 624, 500, 375, 250, 125), sparse attention keeping 20% of tiles.
|
| 15 |
+
`fastvideo_inference.json` holds the schedule, which FastVideo reads automatically.
|
| 16 |
+
- **Text encoder:** NVFP4 Qwen3-VL trimmed to the 50 layers H3 reads.
|
| 17 |
+
- **VAE:** LynnReal lightweight video VAE with Kijai's INT8 weights; H3 audio VAE.
|
|
|
|
|
|
|
| 18 |
|
| 19 |
+
Blog post: [FastH3 on Consumer Hardware](https://haoailab.com/blogs/fasth3-rtx/) 路 Code: [FastVideo](https://github.com/hao-ai-lab/FastVideo)
|
| 20 |
+
|
| 21 |
+
- **Transformer:** 34.9 GiB in BF16 (FP16 AdaLN factors). Source for the quantized releases and for local MLX conversion on Apple Silicon.
|
| 22 |
+
|
| 23 |
+
Full-quality counterpart in the same format: [FastVideo/FastVideo-FastH3-8-Step-V2](https://huggingface.co/FastVideo/FastVideo-FastH3-8-Step-V2).
|