Agnes-Video-2.0-MoE
Agnes-Video-2.0-MoE is an audio-video generation model based on DaVinci-MagiHuman. It accepts text, video, and audio inputs and generates video and audio.
Model Description
Agnes-Video-2.0-MoE uses a 40-layer, single-stream multimodal DiT. The shared feed-forward networks in layers 4–35 are converted to three-expert SparseMoE blocks with Top-1 routing. The checkpoint currently uploaded is the unpruned three-expert SparseMoE model. EENP 60% pruning uses modality-aware neuron selection across video, audio, and text; the pruned checkpoint will be uploaded later.
| Attribute | Description |
|---|---|
| Base model | DaVinci-MagiHuman |
| Model type | Audio-video generation model |
| Input modalities | Text, video, audio |
| Output modalities | Video, audio |
| DiT modification | Layers 4–35; three experts with Top-1 routing |
| Currently uploaded checkpoint | Unpruned SparseMoE |
| Parameters (stored / active) | 28.721B / 15.299B |
| Planned pruned checkpoint | EENP 60%; weights will be uploaded later |
| Planned elastic inference | Quality / Fast widths |
For the planned EENP 60% checkpoint, the nested inference widths are Quality, which retains 65% of intermediate neurons, and Fast, which retains 40% as a high-score subset of Quality. A 32-step run uses 8 Quality steps and 24 Fast steps. These percentages describe routed FFN neuron widths, not the fraction of total model parameters removed. Stored parameters count all routed experts; active parameters count the Top-1 path using the step-averaged schedule. Separately obtained text encoder and VAE weights are not included in the parameter count.
Video Demos
The examples below are text-conditioned generations. Click a preview to open its MP4 video.
Results
Open-Generator Comparison
The comparison uses 32 Koala clips. The “Ours” row is the unpruned SparseMoE checkpoint (28.7B stored / 15.3B active parameters); it is distinct from the EENP 60% checkpoint described above.
Stage Results and Quality–Speed Trade-off
The EENP 60% results below describe the planned pruned checkpoint; its weights will be uploaded in a future update. EENP 60% reduces DiT denoising latency from 26.83 s to 20.92 s (22.0% lower; 1.28× speedup) relative to SparseMoE. VBench 2.0 changes from 0.519 to 0.479. On the paired diagnostic set, FVD changes from 64.267 to 121.012 and CLIP-Sim from 0.9186 to 0.8227, reflecting the quality–speed trade-off.
| Metric | Evaluation setup |
|---|---|
| VBench 2.0 | 448×256; 4 seconds; 32 denoising steps; guidance 2 |
| DiT latency | NVIDIA H100; FP32; batch size 1; mean over 128 five-second generations after one warm-up; excludes VAE decoding |
| Paired diagnostics (FVD / CLIP-Sim) | Fixed 32-clip set; 448×256; 5 seconds; 32 denoising steps; guidance 2; seed 42 |
Inference
Set the component paths in the inference configuration, then run:
export OMNIVAE_CODE_ROOT=/path/to/OmniVAE/vae
export OMNIVAE_RELEASE_ROOT=/path/to/OmniVAE/release
bash scripts/infer.sh \
--config-load-path /path/to/config_moe32_omnivae_init.json \
--ckpt_dir /path/to/agnes_video_2.0_moe_unpruned.pt \
--generate \
--prompt "A small sailboat crossing a calm lake at sunset." \
--seconds 5 \
--num_inference_steps 32 \
--cfg_number 2 \
--save_path_prefix outputs/demo
Model Components
| Component | Source |
|---|---|
| Text encoder | T5Gemma 9B |
| Audio encoder | OmniVAE |
| Video VAE | Wan2.2-TI2V-5B |
License and Third-Party Dependencies
Agnes-Video-2.0-MoE checkpoints contain only Agnes-Video-2.0-MoE DiT/MoE weights. No pretrained base-model weights or third-party model weights are included or redistributed.
Agnes-Video-2.0-MoE code and model weights are provided under the Apache License 2.0, subject to applicable upstream license requirements. See LICENSE and NOTICE.
Agnes-Video-2.0-MoE is developed based on DaVinci-MagiHuman. We acknowledge the original project and retain applicable license notices and attribution.
For inference, users must separately obtain required base-model and auxiliary-model weights from their respective official repositories.
Google T5Gemma is an external text encoder governed by the Gemma Terms of Use, not Apache 2.0. Users must obtain T5Gemma independently and comply with its applicable terms.
Third-party models and dependencies remain subject to their respective licenses. The Apache License 2.0 for Agnes-Video-2.0-MoE does not extend to these external components.
References:








