arXiv:2504.18087cs.CV2025-04中稿 · ACM MM'25被引 15

让说话人脸表情自然且不丢身份,靠解耦与情感协同

Disentangle Identity, Cooperate Emotion: Correlation-Aware Emotional Talking Portrait Generation

  • 分离身份与情绪,用跨模态注意力建模情感
  • 通过可学习情感库捕捉情绪间关联,提升表达连贯性
  • 适合需要高情感真实性和身份一致性的视频生成场景

近期的说话头生成(THG)在扩散模型驱动下实现了出色的口型同步与视觉质量,但现有方法难以在保持说话人身份的同时生成富有情感的表情。我们识别出三大局限:音频情感线索利用不足、情绪表征中存在身份泄露、情绪间关系孤立学习。为此,提出DICE-Talk框架,遵循‘解耦身份与情绪,再协同相似情绪’的理念。首先,设计解耦情绪嵌入器,通过跨模态注意力联合建模音视频情感线索,将情绪表示为与身份无关的高斯分布。其次,引入增强情绪相关性的条件模块,使用可学习的情感库,通过向量量化和注意力聚合显式捕捉情绪间关系。第三,设计情绪判别目标,在扩散过程中通过潜在空间分类强制情感一致性。在MEAD和HDTF数据集上的大量实验表明,本方法在情感准确性上优于现有最优方法,同时保持良好的口型同步性能。定性结果与用户研究进一步验证了其生成身份保持良好、情感丰富且自然联动的说话人脸的能力,可适应未见过的身份。

原文摘要 · Abstract (English)

Recent advances in Talking Head Generation (THG) have achieved impressive lip synchronization and visual quality through diffusion models; yet existing methods struggle to generate emotionally expressive portraits while preserving speaker identity. We identify three critical limitations in current emotional talking head generation: insufficient utilization of audio's inherent emotional cues, identity leakage in emotion representations, and isolated learning of emotion correlations. To address these challenges, we propose a novel framework dubbed as DICE-Talk, following the idea of disentangling identity with emotion, and then cooperating emotions with similar characteristics. First, we develop a disentangled emotion embedder that jointly models audio-visual emotional cues through cross-modal attention, representing emotions as identity-agnostic Gaussian distributions. Second, we introduce a correlation-enhanced emotion conditioning module with learnable Emotion Banks that explicitly capture inter-emotion relationships through vector quantization and attention-based feature aggregation. Third, we design an emotion discrimination objective that enforces affective consistency during the diffusion process through latent-space classification. Extensive experiments on MEAD and HDTF datasets demonstrate our method's superiority, outperforming state-of-the-art approaches in emotion accuracy while maintaining competitive lip-sync performance. Qualitative results and user studies further confirm our method's ability to generate identity-preserving portraits with rich, correlated emotional expressions that naturally adapt to unseen identities.

说话头生成情感表达扩散模型身份保持

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。