Agnes-Video-2.0-MoE

Agnes-Video-2.0-MoE is an audio-video generation model based on DaVinci-MagiHuman. It accepts text, video, and audio inputs and generates video and audio.

Model Description

Agnes-Video-2.0-MoE multimodal DiT and SparseMoE architecture

Agnes-Video-2.0-MoE uses a 40-layer, single-stream multimodal DiT. The shared feed-forward networks in layers 4–35 are converted to three-expert SparseMoE blocks with Top-1 routing. The checkpoint currently uploaded is the unpruned three-expert SparseMoE model. EENP 60% pruning uses modality-aware neuron selection across video, audio, and text; the pruned checkpoint will be uploaded later.

Attribute Description
Base model DaVinci-MagiHuman
Model type Audio-video generation model
Input modalities Text, video, audio
Output modalities Video, audio
DiT modification Layers 4–35; three experts with Top-1 routing
Currently uploaded checkpoint Unpruned SparseMoE
Parameters (stored / active) 28.721B / 15.299B
Planned pruned checkpoint EENP 60%; weights will be uploaded later
Planned elastic inference Quality / Fast widths

For the planned EENP 60% checkpoint, the nested inference widths are Quality, which retains 65% of intermediate neurons, and Fast, which retains 40% as a high-score subset of Quality. A 32-step run uses 8 Quality steps and 24 Fast steps. These percentages describe routed FFN neuron widths, not the fraction of total model parameters removed. Stored parameters count all routed experts; active parameters count the Top-1 path using the step-averaged schedule. Separately obtained text encoder and VAE weights are not included in the parameter count.

Video Demos

The examples below are text-conditioned generations. Click a preview to open its MP4 video.

Orange cat
licking its paw. Preview: Orange cat licking its paw.
Brown-and-white dog
playing on a sunny lawn. Preview: Brown-and-white dog playing on a sunny lawn.
Monkey peeking
over a red apple. Preview: Monkey peeking over a red apple.
White rabbit in a
terracotta flowerpot. Preview: White rabbit sitting inside a terracotta flowerpot.
Golden sunrise over
snow-covered peaks. Preview: Golden sunrise over snow-covered mountain peaks.
Horses grazing beside
a river on grassland. Preview: Horses grazing beside a river on a wide grassland.

Results

Open-Generator Comparison

The comparison uses 32 Koala clips. The “Ours” row is the unpruned SparseMoE checkpoint (28.7B stored / 15.3B active parameters); it is distinct from the EENP 60% checkpoint described above.

Comparison of Agnes-Video-2.0-MoE with open video and audio generators

Stage Results and Quality–Speed Trade-off

The EENP 60% results below describe the planned pruned checkpoint; its weights will be uploaded in a future update. EENP 60% reduces DiT denoising latency from 26.83 s to 20.92 s (22.0% lower; 1.28× speedup) relative to SparseMoE. VBench 2.0 changes from 0.519 to 0.479. On the paired diagnostic set, FVD changes from 64.267 to 121.012 and CLIP-Sim from 0.9186 to 0.8227, reflecting the quality–speed trade-off.

Parameters and scores across Dense, SparseMoE, and EENP 60% stages

Metric Evaluation setup
VBench 2.0 448×256; 4 seconds; 32 denoising steps; guidance 2
DiT latency NVIDIA H100; FP32; batch size 1; mean over 128 five-second generations after one warm-up; excludes VAE decoding
Paired diagnostics (FVD / CLIP-Sim) Fixed 32-clip set; 448×256; 5 seconds; 32 denoising steps; guidance 2; seed 42

Inference

Set the component paths in the inference configuration, then run:

export OMNIVAE_CODE_ROOT=/path/to/OmniVAE/vae
export OMNIVAE_RELEASE_ROOT=/path/to/OmniVAE/release

bash scripts/infer.sh \
  --config-load-path /path/to/config_moe32_omnivae_init.json \
  --ckpt_dir /path/to/agnes_video_2.0_moe_unpruned.pt \
  --generate \
  --prompt "A small sailboat crossing a calm lake at sunset." \
  --seconds 5 \
  --num_inference_steps 32 \
  --cfg_number 2 \
  --save_path_prefix outputs/demo

Model Components

Component Source
Text encoder T5Gemma 9B
Audio encoder OmniVAE
Video VAE Wan2.2-TI2V-5B

License and Third-Party Dependencies

Agnes-Video-2.0-MoE checkpoints contain only Agnes-Video-2.0-MoE DiT/MoE weights. No pretrained base-model weights or third-party model weights are included or redistributed.

Agnes-Video-2.0-MoE code and model weights are provided under the Apache License 2.0, subject to applicable upstream license requirements. See LICENSE and NOTICE.

Agnes-Video-2.0-MoE is developed based on DaVinci-MagiHuman. We acknowledge the original project and retain applicable license notices and attribution.

For inference, users must separately obtain required base-model and auxiliary-model weights from their respective official repositories.

Google T5Gemma is an external text encoder governed by the Gemma Terms of Use, not Apache 2.0. Users must obtain T5Gemma independently and comply with its applicable terms.

Third-party models and dependencies remain subject to their respective licenses. The Apache License 2.0 for Agnes-Video-2.0-MoE does not extend to these external components.

References:

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support