测试视觉模型发现图像逻辑错误的能力,揭示其推理短板。
LADBench: A Benchmark for Logical Fault Detection in Images

- 构建1000+张合成图像,覆盖四大场景的逻辑异常
- 顶级模型准确率仅70.11%,深层提示仍难突破
- 分级提示法可量化模型依赖程度,适合评估智能系统可靠性
大型视觉语言模型在视觉问答和语义定位方面表现优异,但在自主逻辑推理能力方面仍缺乏深入探索。现有异常检测基准多关注视觉误差或直接提问,忽视开放世界部署所需的物理与社会常识。为此,我们提出LAD-bench,一个包含超过1,000张精心设计的合成图像的基准,涵盖住宅、城市、协作和自然四大领域中的逻辑异常。我们进一步提出基于渐进披露的分级提示协议,用于衡量模型在定位和推理逻辑故障时所需显式帮助的程度。对主流基础模型的评估显示其存在显著缺陷:即使最优模型整体准确率也仅为70.11%,表明隐式逻辑故障检测仍未解决。关键的是,模型在获得更深层提示后仍常无法识别异常。通过揭示序列化多模态推理中的局限性,LAD-Bench为提升自主视觉系统的安全性、可靠性和认知一致性提供了严谨框架。数据集与代码:https://huggingface.co/datasets/SahasraK/LADBench
原文摘要 · Abstract (English)
Large Vision Language Models (VLMs) excel at visual question answering and semantic grounding, but their capacity for autonomous logical reasoning remains underexplored. Existing anomaly benchmarks emphasize visual errors or direct prompting rather than the physical and social common sense needed for open-world deployment. To address this, we introduce LAD-bench, a benchmark of more than 1,000 curated synthetic images with logical anomalies across four domains: Residential, Urban, Collaborative, and Nature. We further propose a Tiered Prompting Protocol based on progressive disclosure, which measures how much explicit assistance a model needs to localize and reason about a logical fault. Evaluating leading foundation models reveals substantial weaknesses: even the best achieves only 70.11% overall accuracy, showing that implicit logical fault detection remains unsolved. Crucially, models often fail to identify anomalies even after receiving explicit hints in deeper tiers. By surfacing these limitations in sequential multimodal reasoning, LAD-Bench offers a rigorous framework for advancing the safety, reliability, and cognitive alignment of autonomous visual systems. Dataset and Code: https://huggingface.co/datasets/SahasraK/LADBench
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。