arXiv:2601.16449cs.CVcs.AI2026-01被引 4

新框架+新数据集,让AI更懂多模态情感

Emotion-LLaMAv2 and MMEVerse: A New Framework and Benchmark for Multimodal Emotion Understanding

  • 端到端多视角编码,无需外部人脸检测
  • 13万训练样本,跨18个评测基准提升情感推理能力
  • 适合情感计算、人机交互研究者使用

从多模态信号中理解人类情感是情感计算与人机交互的重大挑战。现有多模态大模型在情感推理方面能力有限,且缺乏高质量、大规模的标注数据集与标准化评估基准。我们提出 Emotion-LLaMAv2 与 MMEVerse 基准,构建端到端情感识别与推理流水线。Emotion-LLaMAv2 引入三项改进:1)端到端多视角编码器,无需外部人脸检测,通过丰富的时空多视图标记捕捉细微情感线索;2)卷积注意力预融合模块,在LLM主干外实现局部与全局特征的同步交互;3)基于感知到认知的课程式指令微调,统一情感识别与自由形式推理。为支持大规模训练与可复现评估,MMEVerse 集成十二个公开情感数据集(如 IEMOCAP、MELD、DFEW、MAFW),经多智能体管道(Qwen2 Audio、Qwen2.5 VL、GPT 4o)重新标注,生成 130,000 条训练片段与 36,000 条测试片段,覆盖 18 个评估基准。

原文摘要 · Abstract (English)

Understanding human emotions from multimodal signals poses a significant challenge in affective computing and human-robot interaction. While multimodal large language models (MLLMs) have excelled in general vision-language tasks, their capabilities in emotional reasoning remain limited. The field currently suffers from a scarcity of large-scale datasets with high-quality, descriptive emotion annotations and lacks standardized benchmarks for evaluation. Our preliminary framework, Emotion-LLaMA, pioneered instruction-tuned multimodal learning for emotion reasoning but was restricted by explicit face detectors, implicit fusion strategies, and low-quality training data with limited scale. To address these limitations, we present Emotion-LLaMAv2 and the MMEVerse benchmark, establishing an end-to-end pipeline together with a standardized evaluation setting for emotion recognition and reasoning. Emotion-LLaMAv2 introduces three key advances. First, an end-to-end multiview encoder eliminates external face detection and captures nuanced emotional cues via richer spatial and temporal multiview tokens. Second, a Conv Attention pre-fusion module is designed to enable simultaneous local and global multimodal feature interactions external to the LLM backbone. Third, a perception-to-cognition curriculum instruction tuning scheme within the LLaMA2 backbone unifies emotion recognition and free-form emotion reasoning. To support large-scale training and reproducible evaluation, MMEVerse aggregates twelve publicly available emotion datasets, including IEMOCAP, MELD, DFEW, and MAFW, into a unified multimodal instruction format. The data are re-annotated via a multi-agent pipeline involving Qwen2 Audio, Qwen2.5 VL, and GPT 4o, producing 130k training clips and 36k testing clips across 18 evaluation benchmarks.

情感理解多模态大模型数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。