MegaAvatar: Controllable Talking Avatar Generation
Paper • 2609.39273 • Published
Junyao Gao*1,2 Sibo Liu*1 Weidong Zhang1 Cairong Zhao2‡ Jun Zhang1‡
1Tencent AIPD 2Tongji University *Equal contribution ‡Corresponding authors
MegaAvatar is a controllable talking avatar generation framework built on top of Wan2.2-TI2V-5B.
Given a reference portrait image, MegaAvatar generates talking avatar videos with:
MegaAvatar supports both SMPL-X-driven and audio-driven generation.
hf download Gaojunyao/MegaAvatar --include "checkpoints/*" --local-dir .
The released checkpoints include:
checkpoints/
├── megaavatar.safetensors
└── audio2smplx/
megaavatar.safetensors: MegaAvatar video generation model.audio2smplx/: audio-to-SMPL-X components.Please refer to the official GitHub repository for installation, inference, and training instructions:
https://github.com/Jeoyal/MegaAvatar
MegaAvatar supports:
Reference Image + SMPL-X + Audio → Talking Avatar Video
Reference Image + Audio → SMPL-X → Talking Avatar Video
This project benefits from FantasyTalking, SpeakerVid-5M-Code, and DiffSynth-Studio.
@article{gao2026megaavatar,
title={MegaAvatar: Controllable Talking Avatar Generation},
author={Gao, Junyao and Liu, Sibo and Zhang, Weidong and Zhao, Cairong and Zhang, Jun},
journal={arXiv preprint arXiv:2609.39273},
year={2026}
}