arXiv:2505.13979cs.CL2025-05被引 2

研究多模态情感识别中模型分歧原因,发现分歧反映真实模糊性。

Mixed Signals: Understanding Model Disagreement in Multimodal Empathy Detection

  • 通过对比单模态与融合模型预测差异,定位冲突根源。
  • 发现单一模态主导信号可能误导融合结果,导致错误判断。
  • 分歧可作诊断信号,帮助提升系统鲁棒性,适合做情感计算研究者参考。

多模态模型在情感识别中至关重要,但当不同模态提供矛盾线索时性能会下降。为理解此类失败,我们分析了单模态与多模态预测不一致的案例。使用微调后的文本、音频和视频模型,以及门控融合模型,发现这些分歧常反映深层语义模糊性,表现为标注者不确定性。分析表明,某一模态的主导信号若缺乏其他模态支持,可能误导融合过程。此外,人类对多模态输入也并非始终受益。这些发现将模型分歧视为识别挑战样本的有效诊断信号,有助于提升情感识别系统的鲁棒性。

原文摘要 · Abstract (English)

Multimodal models play a key role in empathy detection, but their performance can suffer when modalities provide conflicting cues. To understand these failures, we examine cases where unimodal and multimodal predictions diverge. Using fine-tuned models for text, audio, and video, along with a gated fusion model, we find that such disagreements often reflect underlying ambiguity, as evidenced by annotator uncertainty. Our analysis shows that dominant signals in one modality can mislead fusion when unsupported by others. We also observe that humans, like models, do not consistently benefit from multimodal input. These insights position disagreement as a useful diagnostic signal for identifying challenging examples and improving empathy system robustness.

情感识别多模态模型分歧

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。