提出新方法提升多模态讽刺检测的泛化能力
Multi-View Incongruity Learning for Multimodal Sarcasm Detection
- 通过对比学习融合三视图不一致信息增强模型理解
- 在基准数据集上准确率优于现有方法,有效降低误判风险
- 适合关注模型鲁棒性与真实场景泛化能力的研究者
多模态讽刺检测对下游任务至关重要。现有方法易依赖虚假相关性,虽能正确预测但泛化能力差。本文分析其两大成因,提出基于对比学习的多视图不一致融合方法(MICL)。通过文本-图像-情感三视图的不一致驱动学习,并引入大规模数据增强缓解文本模态偏差。构建新测试集SPMSD以评估模型对虚假相关性的抗性。实验表明,MICL在多个基准数据集上表现更优,且显著降低虚假相关性影响。
原文摘要 · Abstract (English)
Multimodal sarcasm detection (MSD) is essential for various downstream tasks. Existing MSD methods tend to rely on spurious correlations. These methods often mistakenly prioritize non-essential features yet still make correct predictions, demonstrating poor generalizability beyond training environments. Regarding this phenomenon, this paper undertakes several initiatives. Firstly, we identify two primary causes that lead to the reliance of spurious correlations. Secondly, we address these challenges by proposing a novel method that integrate Multimodal Incongruities via Contrastive Learning (MICL) for multimodal sarcasm detection. Specifically, we first leverage incongruity to drive multi-view learning from three views: token-patch, entity-object, and sentiment. Then, we introduce extensive data augmentation to mitigate the biased learning of the textual modality. Additionally, we construct a test set, SPMSD, which consists potential spurious correlations to evaluate the the model's generalizability. Experimental results demonstrate the superiority of MICL on benchmark datasets, along with the analyses showcasing MICL's advancement in mitigating the effect of spurious correlation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。