arXiv:2412.17292cs.CVcs.AI2024-12被引 3

让对话系统理解用户表情和语气,生成更共情的回应。

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues

  • 融合语音语调与面部表情,端到端提取情感线索。
  • 在多模态对话中表现优于现有大模型,回应更合情合理。
  • 适合需要高情感智能的交互场景,如客服、陪伴机器人。

在人类交流中,言语与非言语线索共同传递情绪、意图和深层含义。这些非语言信息——如面部表情、眼神接触、语调和音高——是有效互动的基础,能为对话增添情感与情境深度。为重视非语言内容的重要性,我们提出AV-EmoDialog,一种利用用户音视频输入中的言语与非言语信息生成更具响应性与同理心的对话系统。该系统系统性地挖掘音视频对话中的情感线索:从语音中提取话语内容与情感语调,从视觉中分析细微面部表情,并将这些线索融合,以端到端方式生成情感感知的回复。通过大量实验验证,所提出的AV-EmoDialog在生成既情感恰当又上下文合理的回应方面,显著优于现有的多模态大模型。

原文摘要 · Abstract (English)

In human communication, both verbal and non-verbal cues play a crucial role in conveying emotions, intentions, and meaning beyond words alone. These non-linguistic information, such as facial expressions, eye contact, voice tone, and pitch, are fundamental elements of effective interactions, enriching conversations by adding emotional and contextual depth. Recognizing the importance of non-linguistic content in communication, we present AV-EmoDialog, a dialogue system designed to exploit verbal and non-verbal information from users' audio-visual inputs to generate more responsive and empathetic interactions. AV-EmoDialog systematically exploits the emotional cues in audio-visual dialogues; extracting speech content and emotional tones from speech, analyzing fine-grained facial expressions from visuals, and integrating these cues to generate emotionally aware responses in an end-to-end manner. Through extensive experiments, we validate that the proposed AV-EmoDialog outperforms existing multimodal LLMs in generating not only emotionally appropriate but also contextually appropriate responses.

多模态对话情感计算音视频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。