提出新基准与方法,解决多模态情绪识别中模态冲突和缺失问题
EmoMM: Benchmarking and Steering MLLM for Multimodal Emotion Recognition under Conflict and Missingness

- 构建含冲突与缺失数据的多模态情绪识别基准EmoMM
- 发现视频信息在高冗余时被模型忽略(视频贡献崩溃)
- 提出轻量级注意力调节机制CHASE,无需重训练提升可靠性
多模态情绪识别对理解真实交互至关重要。尽管多模态大语言模型(MLLM)在该任务中展现潜力,其在模态冲突与缺失情况下的内部决策机制仍不明确。本文为系统研究此类行为,提出EmoMM基准,包含模态对齐、冲突与缺失三类子集。通过大量实验,我们发现视频贡献崩溃(VCC)现象:由于高令牌冗余和模态偏好,MLLM忽视视频证据。为此,我们提出冲突感知头级注意力引导(CHASE),一种轻量级推理时调节机制,可检测模态冲突并动态调整注意力,有效缓解决策偏差,无需重训练主干模型。实验表明,CHASE在多种设置下均显著提升性能,大幅增强MLLM在复杂情感场景中的可靠性。
原文摘要 · Abstract (English)
Multimodal Emotion Recognition (MER) is critical for interpreting real-world interactions. While Multimodal Large Language Models (MLLM) have shown promise in MER, their internal decision-making mechanisms under modality conflict and missingness remain largely underexplored. In this paper, to systematically investigate these behaviors, we introduce EmoMM, a comprehensive benchmark featuring modality-aligned, conflict, and missing subsets. Through extensive evaluation, we uncover a Video Contribution Collapse (VCC) phenomenon, where MLLM marginalize video evidence due to high token redundancy and modality preferences. To address this, we propose Conflict-aware Head-level Attention Steering (CHASE), a lightweight mechanism that detects modality conflicts and performs inference-time attention steering, effectively mitigating decision bias without retraining the backbone. Experimental results demonstrate that CHASE consistently improves performance across various settings, significantly enhancing the reliability of MLLM in complex affective scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。