arXiv:2607.24430cs.HCcs.CL2026-07中稿 · ACM MM 2026

让语音合成更懂表情,提升对话交互的自然与共情

Let Me Look at You: Advanced Facial Expression Modeling for Conversational Speech Synthesis

论文配图:Let Me Look at You: Advanced Facial Expression Modeling for Conversational Speech Synthesis
图 1 · 摘自论文原文
  • 用视觉编码器将每帧表情转为紧凑标记,结合动作单元监督训练
  • 在1033小时多模态对话数据上,生成语音更自然且表情匹配度高
  • 适合研究情感语音合成、人机交互或多模态模型的开发者

对话式语音合成是人机交互的核心,旨在生成符合语境、富有表现力且具同理心的语音。然而,面部表情蕴含细微而丰富的感情线索,对共情交互至关重要,现有方法常忽视这一模态。此外,缺乏大规模真实对话的音视频同步数据也制约了视觉情感理解的发展。为此,我们提出 FacialTalker,一个基于大语言模型骨干的面部表情感知语音合成框架。为高效编码面部表情,我们设计 AUTokenizer,一种单码本视觉分词器,将每帧表情离散化为紧凑标记,并通过面部动作单元组合进行监督训练。我们进一步引入双通道直接偏好优化(DualDPO),在视觉与语音标记序列上联合施加偏好约束,增强模型对多模态对话中表情与语义的理解。此外,我们构建了 VSDD-1K,一个通过全自动管道从真实网络对话采集的大规模多模态对话数据集,包含超过1033小时同步的说话人视频与语音,其中85%以上的帧含有有效人脸。大量客观与主观实验表明,FacialTalker 在表情感知与语音合成质量上均持续优于强基线,生成的语音更自然、更具表现力且与对话上下文更契合。结果验证了训练策略与数据构建流程的有效性。

原文摘要 · Abstract (English)

Conversational Speech Synthesis is a fundamental component of human-computer interaction, aiming to generate contextually appropriate, expressive, and empathetic speech. However, facial expressions encode subtle and rich affective cues that are crucial for empathetic speech interaction, whereas existing approaches often overlook this important modality. In addition, the lack of large-scale natural conversational datasets with both speech and visual modalities also limits the development of visual affect understanding in conversational settings.To address these limitations, we propose FacialTalker, a facial-expression-aware CSS framework built upon a large language model backbone. To efficiently encode facial expressions, we propose AUTokenizer, a single-codebook visual tokenizer that discretizes each frame-level facial expression into a compact token, trained with supervision from combinations of facial Action Units. We further introduce a dual direct preference optimization (DualDPO) strategy, which extends the DPO by jointly imposing preference constraints on both visual and speech token sequences, to enhance the model's understanding of facial expressions and speech semantics in multimodal conversational contexts. Moreover, we construct VSDD-1K, a large-scale multimodal dialogue dataset collected through a fully automated pipeline from real-world Internet conversations, comprising over 1,033 hours of synchronized speaker videos and speech, with more than 85\% of frames containing valid faces. Extensive objective and subjective experiments demonstrate that FacialTalker consistently outperforms strong baselines in facial-expression perception and speech synthesis quality, generating speech that is more natural, expressive, and better aligned with the conversational context. The results also validate the effectiveness of our training strategy and dataset construction pipeline.

语音合成面部表情多模态大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。