arXiv:2502.16033cs.CLcs.AI2025-02ACL被引 26

新基准测试模型对图文不一致的识别能力,发现主流模型仍易出错。

Multimodal Inconsistency Reasoning (MMIR): A New Benchmark for Multimodal Reasoning Models

  • 构建534个含五类错误的图文不一致样本,评估模型推理能力。
  • 顶尖模型如o1表现显著优于其他模型,开源模型更易出错。
  • 单模态提示效果有限,跨模态推理仍是核心瓶颈。

现有多模态大模型主要在一致的图文输入上训练和测试,难以应对真实场景中布局丰富的不一致内容。为此,我们提出多模态不一致推理(MMIR)基准,用于评估多模态大语言模型在网页、幻灯片、海报等载体中检测和推理语义矛盾的能力。MMIR包含534个挑战性样本,每条数据包含五类高难度推理错误:事实矛盾、身份误指、上下文错配、数量差异、时空不连贯。我们评估了六种前沿多模态模型,结果表明具备专门多模态推理能力的模型(如o1)显著优于其他模型,而开源模型对不一致错误尤为脆弱。详细错误分析显示,模型擅长检测成对不一致,但在复杂布局中单元素内部不一致上表现不佳。探针实验表明,单模态提示(如CoT、SoM)提升有限,揭示了跨模态推理的关键瓶颈。研究强调需发展更先进的多模态推理能力,推动未来关于多模态不一致的研究。

原文摘要 · Abstract (English)

Existing Multimodal Large Language Models (MLLMs) are predominantly trained and tested on consistent visual-textual inputs, leaving open the question of whether they can handle inconsistencies in real-world, layout-rich content. To bridge this gap, we propose the Multimodal Inconsistency Reasoning (MMIR) benchmark to assess MLLMs' ability to detect and reason about semantic mismatches in artifacts such as webpages, presentation slides, and posters. MMIR comprises 534 challenging samples, each containing synthetically injected errors across five reasoning-heavy categories: Factual Contradiction, Identity Misattribution, Contextual Mismatch, Quantitative Discrepancy, and Temporal/Spatial Incoherence. We evaluate six state-of-the-art MLLMs, showing that models with dedicated multimodal reasoning capabilities, such as o1, substantially outperform their counterparts while open-source models remain particularly vulnerable to inconsistency errors. Detailed error analyses further show that models excel in detecting pairwise inconsistencies but struggle with inconsistencies confined to single elements in complex layouts. Probing experiments reveal that single-modality prompting, including Chain-of-Thought (CoT) and Set-of-Mark (SoM) methods, yields marginal gains, revealing a key bottleneck in cross-modal reasoning. Our findings highlight the need for advanced multimodal reasoning and point to future research on multimodal inconsistency.

多模态推理测试基准评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。