评测视觉语言模型对室内3D布局的综合推理能力,填补传统问答评估的空白。
IDEAL-Bench: Indoor Dataset and Evaluation suite for Analyzing 3D Layout reasoning

- 基于真实感室内场景生成1000个可重渲染环境,要求模型从单图推断完整3D布局。
- 最强模型仅得62.1分(满分100),显示当前技术仍严重不足。
- 揭示模型在物体识别与几何推断间的巨大差距,适合研究空间理解的学者使用。
空间问答是评估视觉语言模型(VLMs)空间智能的主要范式,但忽略了另一个重要维度:整体3D布局推理,即从单张图像结构化地预测所有可见物体的位置和范围。为此,我们提出IDEAL-Bench,一个评估套件,要求VLM在涵盖10种房间类型的逼真室内场景中预测结构化的3D布局,评分包含五个数值维度和一种感知渲染对比协议。通过在受控光照与视角下使用语义真实的场景并完全替换资产,IDEAL-Bench超越了CLEVR风格的简单几何体,使任何图像差异仅反映空间推理能力。该基准基于IDEAL-Scenes,一个由1000个可重渲染Blender环境组成的程序化数据集,附带真实布局标注。评估15个主流VLM发现:该任务仍远未解决,最强模型仅达62.1/100;所有模型在物体识别与几何回归间表现出显著不对称,表明当前VLM更擅长描述而非测量;模型排名在部分基准上与问答和原始重建基准不一致,顶级模型共识保持,但中等水平排名大幅变化。这些结果确立IDEAL-Bench作为诊断工具,针对问答评估无法揭示的几何与结构能力,为下一代VLM的空间智能提供更严格的评测路径。
原文摘要 · Abstract (English)
Spatial question answering is the dominant paradigm for evaluating spatial intelligence in Vision-Language Models (VLMs), but it leaves a complementary axis of spatial competence under-evaluated: holistic 3D layout inference, which predicts every visible object's pose and extent from a single image in a structured form. To this end, we introduce IDEAL-Bench, an evaluation suite that requires VLMs to predict structured 3D layouts on photorealistic indoor scenes across 10 room types, scored along five numerical dimensions and a perceptual render-and-compare protocol. By operating on semantically realistic scenes with full asset substitution under controlled lighting and viewpoint, IDEAL-Bench moves beyond CLEVR-style simple geometric primitives so that any image-space discrepancy reflects spatial reasoning alone. The benchmark is built on IDEAL-Scenes, a procedurally generated dataset of 1,000 re-renderable Blender environments with ground-truth layouts. Evaluating 15 prominent VLMs reveals three findings: the task remains substantially unsolved, with the strongest model reaching only 62.1/100 overall; all models exhibit a sharp asymmetry between object recognition and geometric regression, indicating that current VLMs are trained to describe scenes rather than to measure them; model rankings partially diverge from those on QA-based and primitive-reconstruction benchmarks: top-tier consensus holds, but mid-tier rankings shift substantially. Collectively, these findings establish IDEAL-Bench as a diagnostic suite, targeting the geometric and structural competencies that QA-based evaluation cannot surface, and paving the way towards more rigorous evaluation of spatial intelligence in next-generation VLMs. Together, these findings position IDEAL-Bench as a principled diagnostic for whether future VLMs achieve genuine spatial understanding rather than linguistic approximations of it.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。