arXiv:2608.09772cs.CL2026-08

提出新基准,测试大模型是否真懂图文讽刺关系。

PragMatch: Separating Pragmatic Incongruity from Cross-Modal Mismatch in Large Vision-Language Models

论文配图:PragMatch: Separating Pragmatic Incongruity from Cross-Modal Mismatch in Large Vision-Language Models
图 1 · 摘自论文原文
  • 构建3000对图文样本,区分讽刺与表面不匹配。
  • 发现模型依赖文字、OCR和风格线索,易被干扰误导。
  • 适合研究多模态推理、讽刺检测与模型可解释性的学者。

大型视觉语言模型(LVLMs)在多模态基准上表现强劲,但其是否真正理解图像与文本间的关系,还是仅依赖表层相关性(即捷径学习),仍不清楚。这一问题在多模态讽刺检测中尤为关键,因为正确识别依赖于对语用不一致的理解,而非简单的图文不匹配。本文提出PragMatch,一个由MMSD2.0衍生的受控基准,包含3000个图文对,涵盖原始讽刺样本、构造的字面语义对与困难负样本。通过系统性遮蔽分析与针对性信号注入实验,我们发现LVLM预测对词汇、OCR提取内容及风格线索高度敏感;即使图像-文本语义关系不变,注入表面信号也会显著改变模型输出。结果揭示了当前LVLM在语用推理上的局限性,而PragMatch为超越表层对齐的多模态语用推理评估提供了系统化测试平台。

原文摘要 · Abstract (English)

Large Vision-Language Models (LVLMs) have demonstrated strong performance on multimodal benchmarks, yet it remains unclear whether they genuinely reason about relationships between images and text or rely on superficial correlations, known as shortcut learning. This question is particularly important for multimodal sarcasm detection, where successful prediction depends on recognizing pragmatic incongruity rather than treating sarcasm as simple image-text mismatch. We introduce PragMatch, a controlled benchmark of 3,000 image-text pairs derived from MMSD2.0, including original sarcastic examples and constructed literal and hard-negative pairs. We identify influential shortcut cues through systematic masking and evaluate their impact through targeted injection experiments. Our results show that LVLM predictions are sensitive to lexical, OCR-derived and stylistic cues, with injected surface signals causing substantial changes in model predictions despite unchanged underlying image-text relationships. Our findings reveal limitations in current LVLMs while PragMatch provides a systematic testbed for evaluating multimodal pragmatic reasoning beyond surface-level image-text alignment.

多模态讽刺检测语用推理模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。