arXiv:2601.18321cs.MMcs.CL2026-01被引 2

构建细粒度跨模态情感推理框架,提升复杂情境下情绪判断的准确性。

Integrating Fine-Grained Audio-Visual Evidence for Robust Multimodal Emotion Reasoning

  • 采用分步感知与推理机制,分离特征提取与逻辑推断过程。
  • 在60万视频上训练,六维标注体系捕捉音频视觉线索及因果关系。
  • 适合研究多模态情感分析、人机交互与具身智能的开发者使用。

多模态情感分析正从静态分类转向生成式推理。除标签预测外,鲁棒的情感推理需融合面部微表情、语调等细粒度信号,以解码复杂社会情境中的潜在因果关系。然而,当前多模态大语言模型(MLLM)在细粒度感知方面受限于数据稀缺和跨模态融合不足,常出现单模态主导现象,导致在视觉与听觉线索微妙、模糊或矛盾时(如讽刺场景)产生幻觉。为此,我们提出SABER-LLM框架:首先构建SABER数据集,包含60万视频片段,采用新颖的六维标注体系,联合标注视听线索与因果逻辑;其次提出结构化证据分解范式,强制执行“感知-推理”分离,缓解单模态主导问题;并通过一致性感知的直接偏好优化,增强模态间在模糊或冲突条件下的对齐能力。在EMER、EmoBench-M和SABER-Test上的实验表明,SABER-LLM显著优于开源基线,其鲁棒性可媲美闭源模型,能有效解析复杂情感动态。数据集与模型已开源:https://github.com/zxzhao0/SABER-LLM。

原文摘要 · Abstract (English)

Multimodal emotion analysis is shifting from static classification to generative reasoning. Beyond simple label prediction, robust affective reasoning must synthesize fine-grained signals such as facial micro-expressions and prosodic which shifts to decode the latent causality within complex social contexts. However, current Multimodal Large Language Models (MLLMs) face significant limitations in fine-grained perception, primarily due to data scarcity and insufficient cross-modal fusion. As a result, these models often exhibit unimodal dominance which leads to hallucinations in complex multimodal interactions, particularly when visual and acoustic cues are subtle, ambiguous, or even contradictory (e.g., in sarcastic scenery). To address this, we introduce SABER-LLM, a framework designed for robust multimodal reasoning. First, we construct SABER, a large-scale emotion reasoning dataset comprising 600K video clips, annotated with a novel six-dimensional schema that jointly captures audiovisual cues and causal logic. Second, we propose the structured evidence decomposition paradigm, which enforces a "perceive-then-reason" separation between evidence extraction and reasoning to alleviate unimodal dominance. The ability to perceive complex scenes is further reinforced by consistency-aware direct preference optimization, which explicitly encourages alignment among modalities under ambiguous or conflicting perceptual conditions. Experiments on EMER, EmoBench-M, and SABER-Test demonstrate that SABER-LLM significantly outperforms open-source baselines and achieves robustness competitive with closed-source models in decoding complex emotional dynamics. The dataset and model are available at https://github.com/zxzhao0/SABER-LLM.

多模态情感推理细粒度感知大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。