MegaAvatar: Controllable Talking Avatar Generation

Junyao Gao*1,2   Sibo Liu*1   Weidong Zhang1   Cairong Zhao2‡   Jun Zhang1‡

1Tencent AIPD   2Tongji University   *Equal contribution   ‡Corresponding authors

GitHub arXiv Dataset


MegaAvatar is a controllable talking avatar generation framework built on top of Wan2.2-TI2V-5B.

Given a reference portrait image, MegaAvatar generates talking avatar videos with:

  • controllable body and head motion using SMPL-X;
  • speech-synchronized lip motion and facial expressions using audio;
  • identity-preserving facial appearance.

MegaAvatar supports both SMPL-X-driven and audio-driven generation.

Model Download

hf download Gaojunyao/MegaAvatar --include "checkpoints/*" --local-dir .

The released checkpoints include:

checkpoints/
├── megaavatar.safetensors
└── audio2smplx/
  • megaavatar.safetensors: MegaAvatar video generation model.
  • audio2smplx/: audio-to-SMPL-X components.

Usage

Please refer to the official GitHub repository for installation, inference, and training instructions:

https://github.com/Jeoyal/MegaAvatar

MegaAvatar supports:

Reference Image + SMPL-X + Audio → Talking Avatar Video
Reference Image + Audio          → SMPL-X → Talking Avatar Video

Related Resources

Acknowledgements

This project benefits from FantasyTalking, SpeakerVid-5M-Code, and DiffSynth-Studio.

Citation

@article{gao2026megaavatar,
  title={MegaAvatar: Controllable Talking Avatar Generation},
  author={Gao, Junyao and Liu, Sibo and Zhang, Weidong and Zhao, Cairong and Zhang, Jun},
  journal={arXiv preprint arXiv:2609.39273},
  year={2026}
}
Downloads last month
5
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for Gaojunyao/MegaAvatar