MANGO实现多人自然对话的3D虚拟人生成,无需伪3D标签。
MANGO:Natural Multi-speaker 3D Talking Head Generation via 2D-Lifted Enhancement
- 两阶段框架:先用扩散模型建模多说话人语音交互,再用3D高斯渲染提供2D监督
- 在500+身份、50小时数据上实现高保真对话动作生成,超越现有方法
- 适合虚拟会议、数字人交互等需自然双向对话的应用场景
当前音频驱动的3D头像生成主要聚焦单说话人场景,缺乏自然的双向听-说互动。实现流畅的对话行为转换仍是关键挑战。现有3D对话虚拟人方法依赖易出错的伪3D标签,无法捕捉精细面部动态。为此,我们提出新型两阶段框架MANGO,通过纯图像级监督交替训练,有效缓解伪3D标签引入的噪声,从而更贴近真实对话行为。第一阶段采用基于扩散的Transformer与双语音交互模块,从多说话人音频中建模自然3D运动;第二阶段使用快速3D高斯渲染器生成高保真图像,并通过交替训练提供2D级光度监督。此外,我们构建了MANGO-Dialog数据集,包含超过50小时、跨500+身份的对齐2D-3D对话数据。大量实验表明,该方法在两人3D对话动作建模上达到卓越精度与真实感,显著提升音频驱动虚拟人的保真度与可控性。
原文摘要 · Abstract (English)
Current audio-driven 3D head generation methods mainly focus on single-speaker scenarios, lacking natural, bidirectional listen-and-speak interaction. Achieving seamless conversational behavior, where speaking and listening states transition fluidly remains a key challenge. Existing 3D conversational avatar approaches rely on error-prone pseudo-3D labels that fail to capture fine-grained facial dynamics. To address these limitations, we introduce a novel two-stage framework MANGO, which leveraging pure image-level supervision by alternately training to mitigate the noise introduced by pseudo-3D labels, thereby achieving better alignment with real-world conversational behaviors. Specifically, in the first stage, a diffusion-based transformer with a dual-audio interaction module models natural 3D motion from multi-speaker audio. In the second stage, we use a fast 3D Gaussian Renderer to generate high-fidelity images and provide 2D-level photometric supervision for the 3D motions through alternate training. Additionally, we introduce MANGO-Dialog, a high-quality dataset with over 50 hours of aligned 2D-3D conversational data across 500+ identities. Extensive experiments demonstrate that our method achieves exceptional accuracy and realism in modeling two-person 3D dialogue motion, significantly advancing the fidelity and controllability of audio-driven talking heads.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。