arXiv:2511.19199cs.CVcs.AI2025-11被引 2

构建首个跨模态矛盾检测基准,评估模型识破图文冲突能力。

CLASH: A Benchmark for Cross-Modal Contradiction Detection

  • 用控制变量的图文矛盾对构建测试集,区分物体与属性级矛盾。
  • 现有模型在识别跨模态矛盾上表现差,暴露明显模态偏差。
  • 针对性微调可显著提升模型发现图文冲突的能力,适合可信AI研究者。

现实场景中图文矛盾普遍存在,但现有基准多假设输入一致,未能评估跨模态矛盾检测——防止幻觉、保障可靠性的基础能力。我们提出CLASH,一个新型多模态矛盾检测基准,包含与COCO图像配对的矛盾描述,矛盾类型为可控的物体级或属性级。数据集包含针对特定问题的多项选择与开放式问答。基准提供经自动化质量检查筛选的大规模微调集,以及小规模人工验证的诊断集。对主流模型的分析显示其在识别跨模态冲突方面存在严重缺陷,暴露系统性模态偏差和类别特异性弱点。此外,我们实证表明在CLASH上进行针对性微调能显著增强矛盾检测能力。

原文摘要 · Abstract (English)

Contradictory multimodal inputs are common in real-world settings, yet existing benchmarks typically assume input consistency and fail to evaluate cross-modal contradiction detection - a fundamental capability for preventing hallucinations and ensuring reliability. We introduce CLASH, a novel benchmark for multimodal contradiction detection, featuring COCO images paired with contradictory captions containing controlled object-level or attribute-level contradictions. The samples include targeted questions evaluated in both multiple-choice and open-ended formats. The benchmark provides an extensive fine-tuning set filtered through automated quality checks, alongside a smaller human-verified diagnostic set. Our analysis of state-of-the-art models reveals substantial limitations in recognizing cross-modal conflicts, exposing systematic modality biases and category-specific weaknesses. Furthermore, we empirically demonstrate that targeted fine-tuning on CLASH substantially enhances conflict detection capabilities.

多模态矛盾检测基准测试可信AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。