arXiv:2510.25234cs.CVcs.AI2025-10中稿 · ICXR 2025 conferen…

分离语音与表情驱动的3D人脸动画,让虚拟人说话带情绪更自然。

Learning Disentangled Speech- and Expression-Driven Blendshapes for 3D Talking Face Animation

  • 将表情和语音驱动的面部变形建模为线性叠加,实现解耦控制。
  • 在真实数据上训练后,能准确同步口型并自然表达指定情绪。
  • 适合需要高情感表达的虚拟人、数字主播等场景应用。

表情是传达人类情感的基础。随着AI生成内容(AIGC)的快速发展,逼真且富有表现力的3D面部动画日益重要。尽管语音驱动的口型同步技术已有进展,但生成带有情感表达的3D说话人脸仍研究不足。主要障碍在于真实情感3D说话人脸数据集稀缺,采集成本高。为此,我们把面部动画建模为语音与情绪驱动的线性叠加问题。利用包含中性表情的3D说话人脸数据集VOCAset和3D表情序列数据集Florence4D,联合学习由语音和情绪驱动的混合形状(blendshapes)。引入稀疏性约束损失,促进两类混合形状的解耦,同时保留训练数据中固有的跨域二次变形。学习到的混合形状可映射至FLAME模型的表达参数与下颌姿态,用于3D Gaussian avatars的动画生成。定性与定量实验表明,该方法能自然生成具有特定表情的说话人脸,同时保持精准口型同步。感知研究表明,相比现有方法,本方法在情感表达上表现更优,且不牺牲口型同步质量。

原文摘要 · Abstract (English)

Expressions are fundamental to conveying human emotions. With the rapid advancement of AI-generated content (AIGC), realistic and expressive 3D facial animation has become increasingly crucial. Despite recent progress in speech-driven lip-sync for talking-face animation, generating emotionally expressive talking faces remains underexplored. A major obstacle is the scarcity of real emotional 3D talking-face datasets due to the high cost of data capture. To address this, we model facial animation driven by both speech and emotion as a linear additive problem. Leveraging a 3D talking-face dataset with neutral expressions (VOCAset) and a dataset of 3D expression sequences (Florence4D), we jointly learn a set of blendshapes driven by speech and emotion. We introduce a sparsity constraint loss to encourage disentanglement between the two types of blendshapes while allowing the model to capture inherent secondary cross-domain deformations present in the training data. The learned blendshapes can be further mapped to the expression and jaw pose parameters of the FLAME model, enabling the animation of 3D Gaussian avatars. Qualitative and quantitative experiments demonstrate that our method naturally generates talking faces with specified expressions while maintaining accurate lip synchronization. Perceptual studies further show that our approach achieves superior emotional expressivity compared to existing methods, without compromising lip-sync quality.

3D人脸表情生成语音驱动解耦建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。