构建首个覆盖多层级的3D视觉定位基准,揭示现有模型在空间与部件级理解上的严重不足。
From Objects to Anywhere: A Holistic Benchmark for Multi-level Visual Grounding in 3D Scenes
- 设计四层跨度的3D视觉定位数据集,涵盖从人体活动区到物体细部的全域场景
- 顶尖模型在空间级任务上准确率仅30%,部件级约40%,远低于对象级表现
- 揭示当前模型缺乏对3D空间关系和物体构成的精细感知能力,适合场景理解研究者
3D视觉定位在复杂场景中定位物体方面已取得显著进展,但对物体之外的3D场景内容进行指代定位仍属空白。本文提出Anywhere3D-Bench,一个涵盖2,886个指代表达-3D边界框对的综合性3D视觉定位基准,覆盖四个层级:人体活动区域、物体外未占用空间、场景中的单个物体以及物体细部。我们在该基准上评估了多种前沿3D视觉定位方法,以及大语言模型(LLMs)和多模态大语言模型(MLLMs)。实验结果表明,空间级和部件级视觉定位最具挑战性:前者需全面的空间推理能力,如建模3D空间中的距离与相对关系;后者则要求对物体构成的精细感知。即便最佳模型如Google Gemini-2.5-Pro和OpenAI o3,在空间级任务上准确率也仅约30%,部件级约40%,显著低于其在区域级和对象级任务的表现。这些发现凸显当前模型在超越物体语义理解3D场景方面的关键缺陷。
原文摘要 · Abstract (English)
3D visual grounding has made notable progress in localizing objects within complex 3D scenes. However, grounding referring expressions beyond objects in 3D scenes remains unexplored. In this paper, we introduce Anywhere3D-Bench, a holistic 3D visual grounding benchmark consisting of 2,886 referring expression-3D bounding box pairs spanning four different grounding levels: human-activity areas, unoccupied space beyond objects, individual objects in the scene, and fine-grained object parts. We assess a range of state-of-the-art 3D visual grounding methods alongside large language models (LLMs) and multimodal LLMs (MLLMs) on Anywhere3D-Bench. Experimental results reveal that space-level and part-level visual grounding pose the greatest challenges: space-level tasks require a more comprehensive spatial reasoning ability, for example, modeling distances and spatial relations within 3D space, while part-level tasks demand fine-grained perception of object composition. Even the best-performing models, Google Gemini-2.5-Pro and OpenAI o3, achieve just around 30% accuracy on space-level tasks and around 40% on part-level tasks, significantly lower than its performance on area-level and object-level tasks. These findings underscore a critical gap in current models' capacity to understand and reason about 3D scenes beyond object-level semantics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。