构建可无限生成复杂场景的视觉空间推理评测工具
InfiniBench: Infinite Benchmarking for Visual Spatial Reasoning with Customizable Scene Complexity
- 用自然语言生成可控复杂度的3D场景并转为逼真视频
- 在高复杂度场景下生成质量优于现有方法
- 适合研究视觉语言模型在空间推理中的失效模式
现代视觉语言模型需具备应对多样场景复杂度的空间推理能力,但现有评测基准缺乏多样性、可扩展性与完全可定制性,难以隔离分析模型在不同空间条件下的失败模式。为此,本文提出InfiniBench,一个全自动、可定制、用户友好的基准生成器,能参数化控制复杂度,合成理论上无限多的3D场景。其核心创新包括:1)基于大模型的智能框架,迭代优化由自然语言描述生成的程序化场景约束;2)灵活的聚类布局优化器,生成此前程序化方法难以处理的密集杂乱场景;3)任务感知相机轨迹优化方法,确保生成视频覆盖所有物体,适合作为VLM输入。实验表明,InfiniBench在提示保真度和物理合理性方面优于当前最先进的程序化与基于大模型的3D生成方法,尤其在高复杂度场景中表现突出。进一步展示了其在测量、视角转换和时空跟踪等典型空间推理任务上的应用价值。
原文摘要 · Abstract (English)
Modern vision-language models (VLMs) are expected to have abilities of spatial reasoning with diverse scene complexities, but evaluating such abilities is difficult due to the lack of benchmarks that are not only diverse and scalable but also fully customizable. Existing benchmarks offer limited customizability over the scene complexity and are incapable of isolating and analyzing specific VLM failure modes under distinct spatial conditions. To address this gap, instead of individually presenting benchmarks for different scene complexities, in this paper we present InfiniBench, a fully automated, customizable and user-friendly benchmark generator that can synthesize a theoretically infinite variety of 3D scenes with parameterized control on scene complexity. InfiniBench uniquely translates scene descriptions in natural language into photo-realistic videos with complex and physically plausible 3D layouts. This is achieved through three key innovations: 1) a LLM-based agentic framework that iteratively refines procedural scene constraints from scene descriptions; 2) a flexible cluster-based layout optimizer that generates dense and cluttered scenes previously intractable for procedural methods; and 3) a task-aware camera trajectory optimization method that renders scenes into videos with full object coverage as VLM input. Experiments demonstrate that InfiniBench outperforms state-of-the-art procedural and LLM-based 3D generation methods in prompt fidelity and physical plausibility, especially in high-complexity scenarios. We further showcased the usefulness of InfiniBench, by generating benchmarks for representative spatial reasoning tasks including measurement, perspective-taking and spatiotemporal tracking.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。