让AI在未知场景中逐步理解语言指令定位物体
OpenGround: Planning-based Online Perception for Open-World 3D Visual Grounding
- 用任务链规划分解复杂指令,分步引导定位
- 在线感知新物体,比现有方法在新数据上提升17.6%
- 适合开放世界、动态目标场景的智能系统研发
3D视觉定位旨在根据自然语言描述在三维场景中定位物体。现有监督方法泛化能力受限,而近期零样本方法通常依赖预定义的物体查询表(OLT),通过单步推理调用视觉语言模型(VLM)定位目标,难以应对未定义目标和复杂查询。为此,我们提出OpenGround,一种新型零样本开放世界3D视觉定位框架,兼容现有零样本方法。OpenGround融合任务链规划,将查询分解为上下文到目标的子目标计划,实现渐进式定位;并引入上下文引导感知,在任务链指导下在线识别新物体。我们还构建了新数据集OpenTarget,包含超过7000个物体-描述对,用于模拟开放世界评估。大量实验表明,OpenGround在Nr3D上表现竞争力,在ScanRefer上达到当前最优,并在OpenTarget上取得17.6%的显著提升。
原文摘要 · Abstract (English)
3D visual grounding aims to locate objects based on natural language descriptions in 3D scenes. Existing supervised methods are limited by generalization and recent zero-shot methods typically rely on a predefined Object Lookup Table (OLT) to query Visual Language Models (VLMs) for reasoning about object locations via a single step grounding, which limits the applications in scenarios with undefined targets and complex queries. To address these problems, we present OpenGround, a novel zero-shot framework for open-world 3D visual grounding that remains compatible with recent zero-shot methods. OpenGround integrates Task-Chain Planning to decompose a query into a plan of context-to-target sub-goals for progressive grounding, and Context-Guided Perception to perceive novel objects online under context guidance from the task chain. We also propose a new dataset named OpenTarget, which contains over 7000 object-description pairs to mimic open-world evaluation. Extensive experiments demonstrate that OpenGround achieves competitive performance on Nr3D, state-of-the-art on ScanRefer, and delivers a substantial 17.6\% improvement on OpenTarget. Project Page at https://why-102.github.io/openground.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。