arXiv:2510.23299cs.CVcs.MM2025-10被引 4

构建多图讽刺检测新基准,提升真实场景下跨图语义理解能力

MMSD3.0: A Multi-Image Benchmark for Real-World Multimodal Sarcasm Detection

  • 设计跨图序列建模机制,捕捉多图间的潜在关联
  • 在多图数据集上实现当前最佳性能,超越已有方法
  • 适用于社交媒体讽刺分析、跨模态理解研究者

尽管多模态讽刺检测取得进展,现有数据集和方法仍主要聚焦单图场景,忽略了多图间潜在的语义与情感关联。为填补这一空白,我们提出MMSD3.0,一个完全由推特和亚马逊评论中收集的多图样本构成的新基准。同时提出跨图推理模型(CIRM),通过针对性的跨图序列建模捕捉隐含的图像间联系,并引入基于文本-图像对应关系的相关性引导细粒度跨模态融合机制,降低信息整合损失。我们建立了全面且具有代表性的基线,实验表明MMSD3.0能更真实反映现实场景。CIRM在MMSD、MMSD2.0和MMSD3.0上均达领先性能,验证其在单图与多图场景下的有效性。数据集与代码公开于https://github.com/ZHCMOONWIND/MMSD3.0。

原文摘要 · Abstract (English)

Despite progress in multimodal sarcasm detection, existing datasets and methods predominantly focus on single-image scenarios, overlooking potential semantic and affective relations across multiple images. This leaves a gap in modeling cases where sarcasm is triggered by multi-image cues in real-world settings. To bridge this gap, we introduce MMSD3.0, a new benchmark composed entirely of multi-image samples curated from tweets and Amazon reviews. We further propose the Cross-Image Reasoning Model (CIRM), which performs targeted cross-image sequence modeling to capture latent inter-image connections. In addition, we introduce a relevance-guided, fine-grained cross-modal fusion mechanism based on text-image correspondence to reduce information loss during integration. We establish a comprehensive suite of strong and representative baselines and conduct extensive experiments, showing that MMSD3.0 is an effective and reliable benchmark that better reflects real-world conditions. Moreover, CIRM demonstrates state-of-the-art performance across MMSD, MMSD2.0 and MMSD3.0, validating its effectiveness in both single-image and multi-image scenarios. Dataset and code are publicly available at https://github.com/ZHCMOONWIND/MMSD3.0.

多模态讽刺检测跨图理解社交媒体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。