arXiv:2604.07786cs.CVcs.LG2026-04中稿 · CVPR被引 1

通过跨模态情感迁移,让语音驱动人脸视频生成更丰富的情绪表达。

Cross-Modal Emotion Transfer for Emotion Editing in Talking Face Video

论文配图:Cross-Modal Emotion Transfer for Emotion Editing in Talking Face Video
图 1 · 摘自论文原文
  • 用语音和表情特征空间的语义向量,解耦情绪与语言信息。
  • 在MEAD和CREMA-D数据集上提升情感准确率14%。
  • 可生成未见过的复杂情绪(如讽刺),适合影视合成与虚拟人开发。

说话人脸生成作为生成模型的核心应用受到广泛关注。为提升合成视频的表现力与真实感,情感编辑至关重要。现有方法受限于表达灵活性,难以生成扩展性情绪。基于标签的方法使用离散情绪类别,无法捕捉丰富情绪;基于音频的方法虽能利用情感丰富的语音信号,但情绪与语言内容在情感语音中纠缠,难以精准表达目标情绪;基于图像的方法依赖高质量正面参考图,且难以获取扩展情绪(如讽刺)的参考数据。为此,我们提出跨模态情感迁移(C-MET)方法,通过建模语音与视觉特征空间间的情感语义向量,实现基于语音生成面部表情。C-MET结合大规模预训练音频编码器与解耦式表情编码器,学习跨模态情感嵌入差异。在MEAD和CREMA-D数据集上的大量实验表明,该方法相较最先进方法提升情感准确率14%,并能生成具有表现力的说话人脸视频,即使面对未见的扩展情绪。代码、模型权重与演示已公开于https://chanhyeok-choi.github.io/C-MET/。

原文摘要 · Abstract (English)

Talking face generation has gained significant attention as a core application of generative models. To enhance the expressiveness and realism of synthesized videos, emotion editing in talking face video plays a crucial role. However, existing approaches often limit expressive flexibility and struggle to generate extended emotions. Label-based methods represent emotions with discrete categories, which fail to capture a wide range of emotions. Audio-based methods can leverage emotionally rich speech signals - and even benefit from expressive text-to-speech (TTS) synthesis - but they fail to express the target emotions because emotions and linguistic contents are entangled in emotional speeches. Images-based methods, on the other hand, rely on target reference images to guide emotion transfer, yet they require high-quality frontal views and face challenges in acquiring reference data for extended emotions (e.g., sarcasm). To address these limitations, we propose Cross-Modal Emotion Transfer (C-MET), a novel approach that generates facial expressions based on speeches by modeling emotion semantic vectors between speech and visual feature spaces. C-MET leverages a large-scale pretrained audio encoder and a disentangled facial expression encoder to learn emotion semantic vectors that represent the difference between two different emotional embeddings across modalities. Extensive experiments on the MEAD and CREMA-D datasets demonstrate that our method improves emotion accuracy by 14% over state-of-the-art methods, while generating expressive talking face videos - even for unseen extended emotions. Code, checkpoint, and demo are available at https://chanhyeok-choi.github.io/C-MET/

情感生成跨模态人脸视频语音驱动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。