arXiv:2510.21864cs.CVcs.GR2025-10SIGGRAPH被引 6

无需标签,通过语音隐式提取情感与身份特征,实现更自然的面部动画生成。

LSF-Animation: Label-Free Speech-Driven Facial Animation via Implicit Feature Representation

  • 从语音中隐式学习情感特征,不依赖人工标注标签。
  • 在3DMEAD数据集上优于当前最优方法,在情感表达和身份泛化上提升显著。
  • 适合需要高自然度、跨说话人通用性的数字人动画应用。

语音驱动的3D面部动画因其生成富有表现力且时间同步的虚拟人而备受关注。尽管近期研究已开始探索情感感知动画,但大多仍依赖显式的独热编码来表示身份与情感,需给定情感与身份标签,限制了对未见说话人的泛化能力。此外,语音中固有的情感线索常被忽略,影响生成动画的真实感与适应性。本文提出LSF-Animation,一种摒弃显式情感与身份特征表示的新框架。该方法通过语音隐式提取情感信息,并从中性面部网格中捕捉身份特征,从而在无须人工标注的前提下,提升对未见说话人及情感状态的泛化能力。此外,引入分层交互融合模块(HIFB),利用融合令牌整合双变换器特征,更有效地融合情感、运动与身份线索。在3DMEAD数据集上的大量实验表明,本方法在情感表现力、身份泛化与动画真实感方面均超越现有最先进方法。代码将开源:https://github.com/Dogter521/LSF-Animation。

原文摘要 · Abstract (English)

Speech-driven 3D facial animation has attracted increasing interest since its potential to generate expressive and temporally synchronized digital humans. While recent works have begun to explore emotion-aware animation, they still depend on explicit one-hot encodings to represent identity and emotion with given emotion and identity labels, which limits their ability to generalize to unseen speakers. Moreover, the emotional cues inherently present in speech are often neglected, limiting the naturalness and adaptability of generated animations. In this work, we propose LSF-Animation, a novel framework that eliminates the reliance on explicit emotion and identity feature representations. Specifically, LSF-Animation implicitly extracts emotion information from speech and captures the identity features from a neutral facial mesh, enabling improved generalization to unseen speakers and emotional states without requiring manual labels. Furthermore, we introduce a Hierarchical Interaction Fusion Block (HIFB), which employs a fusion token to integrate dual transformer features and more effectively integrate emotional, motion-related and identity-related cues. Extensive experiments conducted on the 3DMEAD dataset demonstrate that our method surpasses recent state-of-the-art approaches in terms of emotional expressiveness, identity generalization, and animation realism. The source code will be released at: https://github.com/Dogter521/LSF-Animation.

语音驱动面部动画隐式表示情感生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。