通过推理重构情感线索,提升多模态情绪识别准确性
Follow the Clues, Frame the Truth: Hybrid-evidential Deductive Reasoning in Open-Vocabulary Multimodal Emotion Recognition
- 设计推理协议,分步提出、验证、决定情感判断
- 在模糊场景下超越基线模型,准确率显著提升
- 适合需要可解释性的情绪分析研究者使用
开放词汇多模态情绪识别(OV-MER)因多模态线索的模糊性而具有挑战性,这些线索常源于未观测的情境动态。尽管多模态大语言模型(MLLMs)具备广泛语义覆盖,但其性能常受限于对主导数据先验的过早依赖,导致忽略跨模态的关键互补情感线索。我们提出HyDRA——一种混合证据演绎推理架构,将推理形式化为“提出-验证-决策”流程。通过分层奖励强化学习,使推理路径与最终任务表现对齐,确保多源线索的合理整合。系统评估验证了设计有效性,HyDRA在模糊或冲突场景中持续优于强基线模型,并生成可解释的诊断性证据链。
原文摘要 · Abstract (English)
Open-Vocabulary Multimodal Emotion Recognition (OV-MER) is inherently challenging due to the ambiguity of equivocal multimodal cues, which often stem from distinct unobserved situational dynamics. While Multimodal Large Language Models (MLLMs) offer extensive semantic coverage, their performance is often bottlenecked by premature commitment to dominant data priors, resulting in suboptimal heuristics that overlook crucial, complementary affective cues across modalities. We argue that effective affective reasoning requires more than surface-level association; it necessitates reconstructing nuanced emotional states by synthesizing multiple evidence-grounded rationales that reconcile these observations from diverse latent perspectives. We introduce HyDRA, a Hybrid-evidential Deductive Reasoning Architecture that formalizes inference as a Propose-Verify-Decide protocol. To internalize this abductive process, we employ reinforcement learning with hierarchical reward shaping, aligning the reasoning trajectories with final task performance to ensure they best reconcile the observed multimodal cues. Systematic evaluations validate our design choices, with HyDRA consistently outperforming strong baselines--especially in ambiguous or conflicting scenarios--while providing interpretable, diagnostic evidence traces.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。