arXiv:2508.10644cs.LG2025-08AAAI被引 1

用信息瓶颈提升多模态讽刺检测,避免模型依赖数据捷径

Conditional Information Bottleneck for Multimodal Fusion: Overcoming Shortcut Learning in Sarcasm Detection

  • 引入条件信息瓶颈,强制模型关注跨模态相关特征
  • 在去捷径数据集MUStARD++$^{R}$上达到最佳性能
  • 适合研究情感识别与多模态融合的学者

多模态讽刺检测需在不同模态间捕捉微妙互补信号并过滤无关信息。现有先进方法常从数据集中学习捷径,而非提取真正的讽刺特征。实验表明,捷径学习会损害模型在真实场景中的泛化能力。通过系统性实验,我们揭示了当前多模态融合策略在讽刺检测中的缺陷,强调有效融合的重要性。为此,我们构建了去除捷径信号的MUStARD++$^{R}$数据集,并提出多模态条件信息瓶颈(MCIB)模型,实现高效多模态融合。实验结果表明,MCIB在不依赖捷径学习的情况下表现最优。

原文摘要 · Abstract (English)

Multimodal sarcasm detection is a complex task that requires distinguishing subtle complementary signals across modalities while filtering out irrelevant information. Many advanced methods rely on learning shortcuts from datasets rather than extracting intended sarcasm-related features. However, our experiments show that shortcut learning impairs the model's generalization in real-world scenarios. Furthermore, we reveal the weaknesses of current modality fusion strategies for multimodal sarcasm detection through systematic experiments, highlighting the necessity of focusing on effective modality fusion for complex emotion recognition. To address these challenges, we construct MUStARD++$^{R}$ by removing shortcut signals from MUStARD++. Then, a Multimodal Conditional Information Bottleneck (MCIB) model is introduced to enable efficient multimodal fusion for sarcasm detection. Experimental results show that the MCIB achieves the best performance without relying on shortcut learning.

讽刺检测多模态融合信息瓶颈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。