融合音视频文本多模态情绪信息,生成更自然的带表情口型同步人脸视频。
SynchroRaMa : Lip-Synchronized and Emotion-Aware Talking Face Generation via Multi-Modal Emotion Embedding
- 用音频、文本和情绪特征联合建模情绪嵌入,提升表达丰富度。
- 引入大模型生成场景描述,支持动态动作与高阶语义变化。
- 在多个基准上超越现有方法,用户评测显示视频更自然流畅。
基于音频驱动的人脸生成受到广泛关注,尤其适用于需要生动自然人机交互的应用。然而,现有情绪感知方法多依赖单一模态(音频或图像)进行情绪嵌入,难以捕捉细微情感线索;且大多仅以单张参考图像为条件,限制了对随时间变化的动作或属性的表达能力。为此,我们提出SynchroRaMa框架,通过融合文本(通过情感分析)、音频(语音情绪识别)及音频衍生的效价-唤醒特征,构建多模态情绪嵌入,实现更具真实感的情绪表达与保真度。为确保自然头部运动与精确口型同步,该框架包含一个音频到运动(A2M)模块,可生成与输入音频对齐的运动帧。此外,利用大语言模型(LLM)生成的场景描述作为额外文本输入,使模型能够捕捉动态行为与高层语义属性。结合视觉与文本双重条件,显著提升了时序一致性与视觉真实性。在多个基准数据集上的定量与定性实验表明,SynchroRaMa优于当前最优方法,在图像质量、表情保留与运动真实感方面均有提升。用户研究进一步证实,相比竞品,SynchroRaMa在整体自然度、运动多样性与视频平滑性上获得更高主观评分。
原文摘要 · Abstract (English)
Audio-driven talking face generation has received growing interest, particularly for applications requiring expressive and natural human-avatar interaction. However, most existing emotion-aware methods rely on a single modality (either audio or image) for emotion embedding, limiting their ability to capture nuanced affective cues. Additionally, most methods condition on a single reference image, restricting the model's ability to represent dynamic changes in actions or attributes across time. To address these issues, we introduce SynchroRaMa, a novel framework that integrates a multi-modal emotion embedding by combining emotional signals from text (via sentiment analysis) and audio (via speech-based emotion recognition and audio-derived valence-arousal features), enabling the generation of talking face videos with richer and more authentic emotional expressiveness and fidelity. To ensure natural head motion and accurate lip synchronization, SynchroRaMa includes an audio-to-motion (A2M) module that generates motion frames aligned with the input audio. Finally, SynchroRaMa incorporates scene descriptions generated by Large Language Model (LLM) as additional textual input, enabling it to capture dynamic actions and high-level semantic attributes. Conditioning the model on both visual and textual cues enhances temporal consistency and visual realism. Quantitative and qualitative experiments on benchmark datasets demonstrate that SynchroRaMa outperforms the state-of-the-art, achieving improvements in image quality, expression preservation, and motion realism. A user study further confirms that SynchroRaMa achieves higher subjective ratings than competing methods in overall naturalness, motion diversity, and video smoothness. Our project page is available at <https://novicemm.github.io/synchrorama>.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。