arXiv:2604.00528cs.CVcs.AI2026-04

用视觉语言模型动态追踪物体,直接从图像流重建3D位置,无需预处理点云。

Think, Act, Build: An Agentic Framework with Vision Language Models for Zero-Shot 3D Visual Grounding

  • 设计动态智能体,将2D视觉理解与3D几何重建解耦
  • 在ScanRefer和Nr3D上超越全监督基线,零样本性能领先
  • 解决多视角覆盖不足问题,适合无标注3D定位任务

3D视觉定位旨在通过自然语言描述定位3D场景中的物体。尽管近期基于视觉语言模型(VLMs)的方法探索了零样本可能性,但通常依赖预处理的3D点云,将定位简化为候选框匹配。为此,我们提出“思考、行动、构建”(TAB)框架,将3D-VG任务重构为直接作用于原始RGB-D流的2D到3D生成式重建范式。具体而言,基于专用3D-VG技能,我们的VLM智能体动态调用视觉工具,在2D帧间跟踪并重建目标。为克服严格语义追踪导致的多视角覆盖不足,引入语义锚定几何扩展机制:先在参考视频片段中锚定目标,再利用多视角几何将其空间位置传播至未观测帧。通过相机参数聚合多视角特征,直接将2D视觉线索映射到3D坐标。此外,我们识别现有基准中的参考歧义和类别错误,并人工修正不准确查询。在ScanRefer和Nr3D上的大量实验表明,该框架完全基于开源模型,显著优于先前零样本方法,甚至超越全监督基线。

原文摘要 · Abstract (English)

3D Visual Grounding (3D-VG) aims to localize objects in 3D scenes via natural language descriptions. While recent advancements leveraging Vision-Language Models (VLMs) have explored zero-shot possibilities, they typically suffer from a static workflow relying on preprocessed 3D point clouds, essentially degrading grounding into proposal matching. To bypass this reliance, our core motivation is to decouple the task: leveraging 2D VLMs to resolve complex spatial semantics, while relying on deterministic multi-view geometry to instantiate the 3D structure. Driven by this insight, we propose "Think, Act, Build (TAB)", a dynamic agentic framework that reformulates 3D-VG tasks as a generative 2D-to-3D reconstruction paradigm operating directly on raw RGB-D streams. Specifically, guided by a specialized 3D-VG skill, our VLM agent dynamically invokes visual tools to track and reconstruct the target across 2D frames. Crucially, to overcome the multi-view coverage deficit caused by strict VLM semantic tracking, we introduce the Semantic-Anchored Geometric Expansion, a mechanism that first anchors the target in a reference video clip and then leverages multi-view geometry to propagate its spatial location across unobserved frames. This enables the agent to "Build" the target's 3D representation by aggregating these multi-view features via camera parameters, directly mapping 2D visual cues to 3D coordinates. Furthermore, to ensure rigorous assessment, we identify flaws such as reference ambiguity and category errors in existing benchmarks and manually refine the incorrect queries. Extensive experiments on ScanRefer and Nr3D demonstrate that our framework, relying entirely on open-source models, significantly outperforms previous zero-shot methods and even surpasses fully supervised baselines.

3D定位视觉语言模型零样本智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。