让3D虚拟人说话更自然:口型、表情、动作全统一控制
3DXTalker: Unifying Identity, Lip Sync, Emotion, and Spatial Dynamics in Expressive 3D Talking Avatars
- 通过2D转3D数据管道和解耦表示,解决身份数据少的问题
- 引入帧级音量与情绪信号,提升口型同步与表情细腻度
- 支持提示词控制头姿风格,适合虚拟主播与交互应用
音频驱动的3D虚拟人生成在虚拟沟通、数字人和互动媒体中日益重要,要求虚拟人保持身份一致性、口型与语音同步、表达情绪并具备自然的空间动态,这些共同构成表达力的综合目标。然而,受限于有限的身份数据、窄范围的音频表征和可控性不足,实现仍具挑战。本文提出3DXTalker,通过数据精炼的身份建模、丰富的音频表示和空间动态可控性,实现可扩展的身份建模,利用2D到3D的数据采集流程和解耦表示缓解数据稀缺问题,提升身份泛化能力。引入帧级振幅和情绪线索,超越标准语音嵌入,确保更优的口型同步和细腻的表情调节。这些线索由基于流匹配的Transformer统一建模,生成连贯的面部动态。此外,3DXTalker能生成自然的头部姿态运动,并支持提示词引导的风格化控制。大量实验表明,3DXTalker在统一框架内实现了口型同步、情感表达与头部动态的融合,显著优于现有方法。
原文摘要 · Abstract (English)
Audio-driven 3D talking avatar generation is increasingly important in virtual communication, digital humans, and interactive media, where avatars must preserve identity, synchronize lip motion with speech, express emotion, and exhibit lifelike spatial dynamics, collectively defining a broader objective of expressivity. However, achieving this remains challenging due to insufficient training data with limited subject identities, narrow audio representations, and restricted explicit controllability. In this paper, we propose 3DXTalker, an expressive 3D talking avatar through data-curated identity modeling, audio-rich representations, and spatial dynamics controllability. 3DXTalker enables scalable identity modeling via 2D-to-3D data curation pipeline and disentangled representations, alleviating data scarcity and improving identity generalization. Then, we introduce frame-wise amplitude and emotional cues beyond standard speech embeddings, ensuring superior lip synchronization and nuanced expression modulation. These cues are unified by a flow-matching-based transformer for coherent facial dynamics. Moreover, 3DXTalker also enables natural head-pose motion generation while supporting stylized control via prompt-based conditioning. Extensive experiments show that 3DXTalker integrates lip synchronization, emotional expression, and head-pose dynamics within a unified framework, achieves superior performance in 3D talking avatar generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。