arXiv:2602.21900cs.SDeess.AS2026-02被引 4

让多模态大模型更懂情绪,准确表达情感回应。

EmoOmni: Bridging Emotional Understanding and Expression in Omni-Modal LLMs

  • 引入情感思维链,从多模态感知逐步推导情绪反应。
  • 在真实对话数据上表现媲美更大模型,情绪表达更精准。
  • 适合需要情感理解的智能客服、虚拟助手场景。

多模态大语言模型(Omni-LLMs)推动了人机交互的发展,实现统一的视听感知与语音响应。然而,现有模型在复杂现实场景中常出现浅层理解与情绪错配的问题。这一问题因思考者-说话者架构通过隐状态间接连接而加剧,导致情绪细节丢失。本文提出EmoOmni框架,实现多模态情感对话的准确理解和表达。核心是情感思维链(E-CoT),强制从细粒度多模态感知推理至文本响应;同时将E-CoT显式作为高层情感指令指导说话者,提升情感表达准确性。配套构建了真实标注对话数据集EmoOmniPipe,并建立评估基准EmoOmniEval,支持系统性评测。实验表明,EmoOmni-7B在相同说话者下性能媲美Qwen3Omni-30B-A3B-Thinking。

原文摘要 · Abstract (English)

The evolution of Omni-Modal Large Language Models~(Omni-LLMs) has revolutionized human--computer interaction, enabling unified audio-visual perception and speech response. However, existing Omni-LLMs struggle with complex real-world scenarios, often leading to superficial understanding and contextually mismatched emotional responses. This issue is further intensified by Omni-LLM's Thinker-Talker architectures, which are implicitly connected through hidden states, leading to the loss of emotional details. In this work, we present EmoOmni, a unified framework for accurate understanding and expression in multimodal emotional dialogue. At its core, we introduce the emotional Chain-of-Thought~(E-CoT), which enforces a reasoning from fine-grained multimodal perception to textual response. Moreover, we explicitly treat E-CoT as high-level emotional instructions that guide the talker, enabling accurate emotional expression. Complementing the model, we construct EmoOmniPipe to obtain the real-world annotated dialogue data and establish a benchmark, EmoOmniEval, to facilitate systematic assessment of multimodal emotional dialogue task. Experiments show that EmoOmni-7B achieves comparable performance with Qwen3Omni-30B-A3B-Thinking under the same talker.

多模态情感理解大模型对话系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。