让视觉语言模型通过主动建模3D场景来检验真实理解能力
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering
- 用编程与渲染工具主动重建输入图像的3D结构
- 发现当前模型在视觉精度上存在明显短板
- 适合研究具身智能与多模态生成的学者
视觉语言模型擅长描述任务,但其对场景的真实理解仍不确定。我们提出IR3D-Bench,一个基于分析-合成范式的基准测试,要求视觉语言代理(VLAs)主动使用编程和渲染工具,重建输入图像背后的3D结构,实现通过工具使用的智能逆向渲染。该‘通过创造理解’的方法考察了VLAs的工具使用生成能力,超越传统场景理解评测中的描述或对话能力。我们提供涵盖几何精度、空间关系、外观属性及整体合理性的一整套评估指标。对多种前沿VLM的初步实验表明,当前模型在视觉精度方面存在显著不足,而非基础工具使用能力。IR3D-Bench(含数据与评估协议)已发布,以推动具身智能中工具使用型视觉语言模型的发展,迈向真正的场景理解。
原文摘要 · Abstract (English)
Vision-language models (VLMs) excel at descriptive tasks, but whether they truly understand scenes from visual observations remains uncertain. We introduce IR3D-Bench, a benchmark challenging VLMs to demonstrate understanding through active creation rather than passive recognition. Grounded in the analysis-by-synthesis paradigm, IR3D-Bench tasks Vision-Language Agents (VLAs) with actively using programming and rendering tools to recreate the underlying 3D structure of an input image, achieving agentic inverse rendering through tool use. This "understanding-by-creating" approach probes the tool-using generative capacity of VLAs, moving beyond the descriptive or conversational capacity measured by traditional scene understanding benchmarks. We provide a comprehensive suite of metrics to evaluate geometric accuracy, spatial relations, appearance attributes, and overall plausibility. Initial experiments on agentic inverse rendering powered by various state-of-the-art VLMs highlight current limitations, particularly in visual precision rather than basic tool usage. IR3D-Bench, including data and evaluation protocols, is released to facilitate systematic study and development of tool-using VLAs towards genuine scene understanding by creating.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。