arXiv:2605.04877cs.MMcs.HC2026-05

提出双路径框架,智能决定多模态情绪识别中何时融合、何时舍弃模态。

To Fuse or to Drop? Dual-Path Learning for Resolving Modality Conflicts in Multimodal Emotion Recognition

论文配图:To Fuse or to Drop? Dual-Path Learning for Resolving Modality Conflicts in Multimodal Emotion Recognition
图 1 · 摘自论文原文
  • 设计双路径机制:一路校准模态,一路决策是否融合。
  • 在五个基准上超越现有方法,尤其在矛盾场景下表现更优。
  • 适合需要鲁棒多模态情绪识别的对话系统与视频分析场景。

多模态情绪识别(MER)虽能融合文本、音频和视觉信息,但当模态间存在冲突时,标准融合方法常失效。关键在于冲突可分两类:良性冲突源于信息缺失或模糊,可通过跨模态校准缓解;严重冲突来自内在矛盾(如反语)或误导信号,强行融合反而放大错误。为此,我们提出双路径冲突解决框架(DCR),自动学习何时融合、何时丢弃模态。路径一(情感融合蒸馏器,AFD)通过时间加权类别证据,从音频/视觉教师模型向文本学生模型进行反向蒸馏,增强表征级校准,提升对齐有益时的融合效果。路径二(情感辨识代理,ADA)将MER建模为上下文老虎机问题,基于双视角状态和校准感知奖励,在融合与单模态预测间做出决策,无需每模态可靠性标签即可实现不可调和冲突下的决策级仲裁。通过结合软校准与硬决策,DCR既能协调可对齐的冲突,又能避开有害融合。在覆盖对话级与片段级的五个基准上,DCR持续优于或媲美主流基线。消融实验、特定冲突子集评估与模态选择分析验证了AFD与ADA的互补性及其对鲁棒冲突感知情绪识别的协同增益。

原文摘要 · Abstract (English)

Multimodal emotion recognition (MER) benefits from combining text, audio, and vision, yet standard fusion often fails when modalities conflict. Crucially, conflicts differ in resolvability: benign conflicts stem from missing, weak, or ambiguous cues and can be mitigated by cross-modal calibration, while severe conflicts arise from intrinsically contradictory (e.g., sarcasm) or misleading signals, for which forced fusion may amplify errors. Recognizing this, we propose Dual-Path Conflict Resolution (DCR), a unified framework that learns when to fuse and when to drop modalities. Path I (Affective Fusion Distiller, AFD) performs reverse distillation from audio/visual teachers to a textual student using temporally weighted class evidence, thereby enhancing representation-level calibration and improving fusion when alignment is beneficial. Path II (Affective Discernment Agent, ADA) formulates MER as a contextual bandit that selects among fusion and unimodal predictions based on a dual-view state and a calibration-aware reward, enabling decision-level arbitration under irreconcilable conflicts without requiring per-modality reliability labels. By taking into account the full multimodal context and coupling soft calibration with hard arbitration, DCR reconciles conflicts that can be aligned while bypassing misleading modalities when fusion is harmful. Across five benchmarks covering both dialogue-level and clip-level MER, DCR consistently outperforms competitive baselines or achieves highly competitive results. Further ablations, conflict-specific subset evaluation, and modality-selection analysis verify that AFD and ADA are complementary and jointly improve robust conflict-aware emotion recognition.

多模态情绪识别冲突解决双路径

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。