提出CRISP诊断视觉空间智能,区分感知与推理能力
From Hallucination to Grounding: Diagnosing Visual Spatial Intelligence via CRISP

- 用3D场景图和预言干预解耦感知与推理瓶颈
- 发现大模型感知不准但推理强,开源模型缺乏多跳推理
- 适合关注多模态对齐与模型诊断的研究者
当前视觉语言模型评估常混淆语言先验与真实空间推理。为此,我们提出CRISP——一种基于结构化诊断的评估范式,通过一致性(隐式感知与显式推理的对齐)来衡量视觉空间智能。不同于传统黑箱问答,CRISP利用度量3D场景图与预言干预协议,分离潜在推理能力与感知瓶颈。该细粒度诊断揭示出系统性的感知-推理断层:尽管专有模型具备强大的潜在推理引擎,却存在度量估计不准及未能利用隐式结构表征的问题;而开源模型则根本受限于缺乏多跳组合推理能力。通过将重点从依赖语言先验‘猜对’转向真正‘感知、验证、推理’,CRISP为超越端到端后训练的多模态对齐提供了严谨路径。代码与数据集见https://github.com/iiyamayuki/CRISP-Bench。
原文摘要 · Abstract (English)
Current VLM evaluations often conflate language priors with genuine spatial reasoning. To address this, we introduce CRISP, a novel structural-diagnostic evaluation paradigm that assesses visual spatial intelligence through consistency, the alignment between implicit perception and explicit reasoning. Unlike traditional black-box QA, CRISP utilizes metric 3D Scene Graphs and an oracle intervention protocol to decouple latent reasoning capabilities from perceptual bottlenecks. This granular diagnosis uncovers a systematic perception-reasoning disconnect. Crucially, we reveal that while proprietary models possess robust latent reasoning engines, they suffer from inaccurate metric estimation and a critical failure to leverage their implicit structural representations. Conversely, open-source models remain fundamentally bottlenecked by their lack of multi-hop compositional reasoning. By shifting the focus from merely ``guessing correctly'' via language priors to genuinely ``perceiving, verifying, and reasoning,'' CRISP offers a rigorous roadmap for multimodal alignment beyond end-to-end post-training. The code and dataset are available at https://github.com/iiyamayuki/CRISP-Bench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。