测试视觉语言模型对物体移动的假设后果推理能力,发现模型表现远低于人类。
MindEdit-Bench: Benchmarking Object-Level Counterfactual Spatial Reasoning in VLMs from In-the-Wild Photos

- 构建六类空间推理任务,聚焦物体级反事实推理。
- 模型平均准确率仅8%-31%,人类达81%-97%。
- 适合评估模型在真实场景下的空间想象力与容错能力。
现有视觉语言模型(VLMs)基准多测试可观测空间推理:模型描述输入中已存在的关系。现有假设性任务通常改变观察者而固定场景。能否让VLM预测假设移动或旋转物体的后果?我们提出MindEdit-Bench,一个基于三张智能手机拍摄的室内场景照片组成的基准,通过自动化的野外3D场景图提取流程构建。包含六项任务:四项测试感知与视角变换;两项新任务——L4(空间编辑)和L5(跨视角可见性编辑),要求模型推断输入图像中不存在的正确答案。每道题提供8-24个结构化选项,支持答案字母级别的错误诊断。基准涵盖120个未公开的私有室内场景,降低公共数据预训练重叠风险。在1,003道人工验证问题上,15个VLM的平均准确率为8%-31%,人类多数投票准确率达81%-97%。人类与最优模型差距达53个百分点,每项任务均至少相差39个百分点。结构化答案空间揭示非均匀失败模式,包括相机深度轴推理较弱及在复杂可见性编辑任务中的退化行为。
原文摘要 · Abstract (English)
Benchmarks for vision-language models (VLMs) mostly test observational spatial reasoning: models describe relations already visible in the input. Existing what-if tasks typically vary the observer while keeping the scene fixed. Can VLMs instead predict the consequences of hypothetically moving or rotating an object? We introduce MindEdit-Bench, a benchmark of six spatial reasoning tasks built from three-photo smartphone triplets of newly captured indoor scenes via an automatic in-the-wild 3D scene-graph extraction pipeline. Four tasks probe perception and perspective transformation over observed structure; two new tasks, L4 (spatial editing) and L5 (cross-view visibility editing), probe object-level counterfactual reasoning, where correct answers are absent from all input images. Each question provides 8-24 structured answer choices, enabling answer-letter-level diagnosis of spatial and fallback errors. The benchmark covers 120 private indoor scenes not drawn from public datasets, reducing public-data pretraining-overlap risk. Across 15 VLMs on 1,003 human-verified questions, task-wise mean VLM accuracy is only 8%-31%, versus 81%-97% human majority-vote accuracy. The pooled human--best-VLM gap is 53 pp, with at least 39 pp on every task. The structured answer space further reveals non-uniform failures, including weaker camera-depth-axis inference and fallback behavior on difficult visibility-editing cases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。