动态融合+反讽感知对比正则,提升多模态反讽检测准确率
Dynamic Gated Cross-Modal Fusion with Sarcastic-aware Contrastive Regularization for Multimodal Sarcasm Detection

- 按实例动态调节文本与视觉贡献,自适应融合跨模态特征
- 在MMSD和MMSD2.0上性能超越强基线,提升显著
- 适合需要精准识别反讽意图的多模态应用如社交内容分析
多模态反讽检测旨在从多源内容中识别反讽意图,其核心在于语义与上下文线索间的不一致。现有方法常采用固定融合策略,将反讽视为一般跨模态不匹配,难以捕捉细微反讽信号及实例特异性模态交互。为此,本文提出一种新型多模态反讽检测框架(MSD),结合动态门控跨模态融合与反讽感知对比正则化(SaCR)。通过双向门控交互模块实现跨模态特征筛选,并在实例层面自适应校准文本与视觉贡献;动态融合门进一步平衡模态重要性,生成更鲁棒的多模态表示。同时,SaCR作为标签感知的对比正则化目标,促使非反讽样本保持语义一致,而抑制反讽样本中的误导性一致性。框架采用端到端多目标学习策略,联合优化多模态分类与辅助单模态监督。在MMSD与MMSD2.0数据集上的大量实验表明,所提方法持续优于强基线。
原文摘要 · Abstract (English)
Multimodal sarcasm detection aims to identify sarcastic intent from multimodal content, where inconsistencies between literal meaning and contextual cues often signal irony. This task has attracted increasing research attention. However, accurate detection remains challenging due to instance-dependent modality contributions and misleading semantic consistency, where surface-level alignment masks underlying contradictory intent. Existing methods often rely on fixed fusion strategies and treat sarcasm as generic cross-modal mismatch, limiting their ability to capture subtle sarcasm cues and instance-specific modality interactions. To address these challenges, we propose a novel MSD framework that integrates Dynamic Gated Cross-Modal Fusion with Sarcastic-aware Contrastive Regularization (SaCR). Specifically, a bidirectional gated interaction module performs cross-modal feature filtering and adaptively calibrates textual and visual contributions at the instance level. A dynamic fusion gate further balances modality importance to generate more robust multimodal representations. Furthermore, SaCR is introduced as a label-aware contrastive regularization objective that encourages semantic consistency for non-sarcastic samples while suppressing misleading consistency in sarcastic cases. The proposed framework is trained end-to-end with a multi-objective learning strategy that jointly optimizes multimodal classification and auxiliary unimodal supervision. Extensive experiments on MMSD and MMSD2.0 demonstrate that the proposed method consistently outperforms strong baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。