arXiv:2505.20511cs.CL2025-05EMNLP综述被引 49

融合文本、语音、视觉多模态信息,提升对话情绪识别准确率

Multimodal Emotion Recognition in Conversations: A Survey of Methods, Trends, Challenges and Prospects

  • 整合文本、语音、视觉三模态信息进行情绪判断
  • 提出系统性框架梳理现有方法与评估策略
  • 适合情感计算、人机交互领域研究者参考

尽管基于文本的情绪识别已取得显著进展,但现实对话系统往往需要超越单一模态的更细腻情感理解。多模态对话情绪识别(MERC)因此成为提升人机交互自然性与情感理解力的关键方向。其目标是通过融合文本、语音和视觉信号等多源信息,实现更精准的情绪识别。本文系统综述了MERC的研究动机、核心任务、代表性方法及评估方式,进一步分析了最新趋势,指出关键挑战,并展望未来发展方向。随着对情感智能系统兴趣的增长,本综述为推进MERC研究提供了及时指导。

原文摘要 · Abstract (English)

While text-based emotion recognition methods have achieved notable success, real-world dialogue systems often demand a more nuanced emotional understanding than any single modality can offer. Multimodal Emotion Recognition in Conversations (MERC) has thus emerged as a crucial direction for enhancing the naturalness and emotional understanding of human-computer interaction. Its goal is to accurately recognize emotions by integrating information from various modalities such as text, speech, and visual signals. This survey offers a systematic overview of MERC, including its motivations, core tasks, representative methods, and evaluation strategies. We further examine recent trends, highlight key challenges, and outline future directions. As interest in emotionally intelligent systems grows, this survey provides timely guidance for advancing MERC research.

情绪识别多模态对话系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。