用生成的客观描述检测图文反差,提升讽刺识别准确率
GDCNet: Generative Discrepancy Comparison Network for Multimodal Sarcasm Detection
- 用多模态大模型生成客观图像描述作语义锚点
- 融合语义与情感差异特征,实现跨模态冲突建模
- 在MMSD2.0上达到新最优,适合跨模态讽刺检测任务
多模态讽刺检测(MSD)旨在通过建模跨模态语义不一致来识别图文对中的讽刺。现有方法常依赖跨模态嵌入错位检测不一致,但在视觉与文本关联松散或语义间接时表现不佳。尽管近期方法利用大语言模型(LLMs)生成讽刺线索,但其生成内容的多样性和主观性常引入噪声。为此,我们提出生成式差异对比网络(GDCNet),通过多模态大模型(MLLMs)生成的事实性、描述性图像标题作为稳定语义锚点,捕捉生成描述与原始文本间的语义与情感差异,并结合视觉-文本保真度测量,构建跨模态冲突特征。这些差异特征通过门控模块与视觉、文本表征融合,自适应调节模态贡献。在多个MSD基准上的大量实验表明,GDCNet在准确性与鲁棒性方面均表现优异,在MMSD2.0上取得新的最优结果。
原文摘要 · Abstract (English)
Multimodal sarcasm detection (MSD) aims to identify sarcasm within image-text pairs by modeling semantic incongruities across modalities. Existing methods often exploit cross-modal embedding misalignment to detect inconsistency but struggle when visual and textual content are loosely related or semantically indirect. While recent approaches leverage large language models (LLMs) to generate sarcastic cues, the inherent diversity and subjectivity of these generations often introduce noise. To address these limitations, we propose the Generative Discrepancy Comparison Network (GDCNet). This framework captures cross-modal conflicts by utilizing descriptive, factually grounded image captions generated by Multimodal LLMs (MLLMs) as stable semantic anchors. Specifically, GDCNet computes semantic and sentiment discrepancies between the generated objective description and the original text, alongside measuring visual-textual fidelity. These discrepancy features are then fused with visual and textual representations via a gated module to adaptively balance modality contributions. Extensive experiments on MSD benchmarks demonstrate GDCNet's superior accuracy and robustness, establishing a new state-of-the-art on the MMSD2.0 benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。