arXiv:2511.21717cs.CLcs.CV2025-11AAAI被引 2

诊断多模态模型在真实冲突场景下的推理短板,揭示其在复杂推理中的普遍失效。

CrossCheck-Bench: Diagnosing Compositional Failures in Multimodal Conflict Resolution

  • 构建分层任务框架,涵盖感知、融合到逻辑推理三类能力,评估跨模态矛盾检测。
  • 15,000个含人工注入矛盾的真实世界数据对中,模型性能随推理复杂度显著下降。
  • 符号推理与视觉理解结合的方法优于传统提示策略,适合研究多模态推理瓶颈的学者。

多模态大语言模型主要在对齐的图文数据上训练和评估,对其在真实世界不一致情况下的检测与解决能力研究不足。开放域应用中,视觉与文本线索常存在冲突,要求模型进行超越表面匹配的结构化推理。我们提出 CrossCheck-Bench,一个用于诊断多模态输入中矛盾检测的基准测试。该基准采用分层任务框架,覆盖三个推理复杂度层级,并定义了七种解决跨模态不一致所必需的原子能力。CrossCheck-Bench 包含 15,000 个来自真实世界实体的问答对,其中矛盾为人工注入。数据集通过超过 450 小时的专家标注流程构建,确保语义有效性及感知、融合、推理层面的难度均衡。我们评估了 13 个前沿视觉-语言模型,发现随着任务从感知匹配转向逻辑矛盾检测,性能持续下降。多数模型在孤立实体识别上表现良好,但在需综合多个线索进行冲突推理时失败。能力级分析显示,多步推理或基于规则验证的任务中技能获取不均。额外探查表明,传统的链式思维(Chain-of-Thought)和标记集合(Set-of-Mark)提示策略仅带来微弱提升。相比之下,将符号推理与具身视觉处理相结合的方法表现出更稳定的改进。结果凸显多模态推理中的持续瓶颈,并指明构建鲁棒跨模态验证模型的新方向。

原文摘要 · Abstract (English)

Multimodal Large Language Models are primarily trained and evaluated on aligned image-text pairs, which leaves their ability to detect and resolve real-world inconsistencies largely unexplored. In open-domain applications visual and textual cues often conflict, requiring models to perform structured reasoning beyond surface-level alignment. We introduce CrossCheck-Bench, a diagnostic benchmark for evaluating contradiction detection in multimodal inputs. The benchmark adopts a hierarchical task framework covering three levels of reasoning complexity and defines seven atomic capabilities essential for resolving cross-modal inconsistencies. CrossCheck-Bench includes 15k question-answer pairs sourced from real-world artifacts with synthetically injected contradictions. The dataset is constructed through a multi-stage annotation pipeline involving more than 450 expert hours to ensure semantic validity and calibrated difficulty across perception, integration, and reasoning. We evaluate 13 state-of-the-art vision-language models and observe a consistent performance drop as tasks shift from perceptual matching to logical contradiction detection. Most models perform well on isolated entity recognition but fail when multiple clues must be synthesized for conflict reasoning. Capability-level analysis further reveals uneven skill acquisition, especially in tasks requiring multi-step inference or rule-based validation. Additional probing shows that conventional prompting strategies such as Chain-of-Thought and Set-of-Mark yield only marginal gains. By contrast, methods that interleave symbolic reasoning with grounded visual processing achieve more stable improvements. These results highlight a persistent bottleneck in multimodal reasoning and suggest new directions for building models capable of robust cross-modal verification.

多模态推理瓶颈矛盾检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。