综述多模态大模型在情感识别与推理中的进展,填补领域系统性回顾空白。
Multimodal Large Language Models Meet Multimodal Emotion Recognition and Reasoning: A Survey
- 梳理多模态大模型在跨模态情感理解中的架构与融合机制
- 汇总主流数据集与性能基准,揭示当前模型在复杂场景下的表现差异
- 适合关注情感计算、多模态建模的研究者参考
近年来,大语言模型(LLMs)推动了语言理解的显著进步,标志着向通用人工智能(AGI)迈进的重要一步。随着对更高层次语义和跨模态融合需求的增长,多模态大语言模型(MLLMs)应运而生,整合文本、视觉、音频等多种信息源,以增强复杂场景下的建模与推理能力。在科学智能(AI for Science)领域,多模态情感识别与推理已成为快速发展的前沿方向。尽管LLMs和MLLMs在此领域已取得显著进展,但该领域仍缺乏系统性的综述。为填补这一空白,本文全面回顾了用于情感识别与推理的LLMs与MLLMs,涵盖模型架构、数据集与性能基准。进一步指出了关键挑战,并展望未来研究方向,旨在为研究人员提供权威参考与实用洞见。据我们所知,这是首个系统性综述多模态大模型与多模态情感识别及推理交叉领域的论文。现有方法的总结详见我们的GitHub:https://github.com/yuntaoshou/Awesome-Emotion-Reasoning。
原文摘要 · Abstract (English)
In recent years, large language models (LLMs) have driven major advances in language understanding, marking a significant step toward artificial general intelligence (AGI). With increasing demands for higher-level semantics and cross-modal fusion, multimodal large language models (MLLMs) have emerged, integrating diverse information sources (e.g., text, vision, and audio) to enhance modeling and reasoning in complex scenarios. In AI for Science, multimodal emotion recognition and reasoning has become a rapidly growing frontier. While LLMs and MLLMs have achieved notable progress in this area, the field still lacks a systematic review that consolidates recent developments. To address this gap, this paper provides a comprehensive survey of LLMs and MLLMs for emotion recognition and reasoning, covering model architectures, datasets, and performance benchmarks. We further highlight key challenges and outline future research directions, aiming to offer researchers both an authoritative reference and practical insights for advancing this domain. To the best of our knowledge, this paper is the first attempt to comprehensively survey the intersection of MLLMs with multimodal emotion recognition and reasoning. The summary of existing methods mentioned is in our Github: \href{https://github.com/yuntaoshou/Awesome-Emotion-Reasoning}{https://github.com/yuntaoshou/Awesome-Emotion-Reasoning}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。