arXiv:2603.09506cs.CVcs.RO2026-03中稿 · CVPR被引 2

用上下文引导探索与视角感知空间推理,精准导航到3D场景中的目标实例。

Context-Nav: Context-Driven Exploration and Viewpoint-Aware 3D Spatial Reasoning for Instance Navigation

  • 通过全局文本-图像对齐生成价值图,指导智能体朝符合完整描述的区域探索。
  • 在候选目标处进行视角感知的空间关系验证,仅当至少一个视角满足关系才接受。
  • 无需任务微调,适用于复杂3D场景中细粒度实例定位,适合机器人导航研究者。

文本目标实例导航(TGIN)要求智能体将一段自由描述转化为动作,从同类别干扰物中找到正确的目标实例。我们提出Context-Nav,将长篇上下文描述从局部匹配线索提升为全局探索先验,并通过3D空间推理验证候选对象。首先,计算密集的文本-图像对齐以生成价值图,对前缘区域进行排序,引导探索向与完整描述一致的区域,而非早期检测结果。其次,在观察候选对象后,执行视角感知的关系检查:智能体采样合理的观察者位姿,对齐局部坐标系,并仅当至少一个视角能满足空间关系时才接受目标。该流程无需任务特定训练或微调,在InstanceNav和CoIN-Bench上达到当前最优性能。消融实验表明:(i) 将完整描述编码进价值图可避免无效运动;(ii) 显式的视角感知3D验证能防止语义合理但错误的停靠。这表明基于几何的空间推理是细粒度实例消歧的一种可扩展替代方案,优于依赖大量策略训练或人工交互的方法。

原文摘要 · Abstract (English)

Text-goal instance navigation (TGIN) asks an agent to resolve a single, free-form description into actions that reach the correct object instance among same-category distractors. We present \textit{Context-Nav}, which elevates long, contextual captions from a local matching cue to a global exploration prior and verifies candidates through 3D spatial reasoning. First, we compute dense text-image alignments for a value map that ranks frontiers -- guiding exploration toward regions consistent with the entire description rather than early detections. Second, upon observing a candidate, we perform a viewpoint-aware relation check: the agent samples plausible observer poses, aligns local frames, and accepts a target only if the spatial relations can be satisfied from at least one viewpoint. The pipeline requires no task-specific training or fine-tuning; we attain state-of-the-art performance on InstanceNav and CoIN-Bench. Ablations show that (i) encoding full captions into the value map avoids wasted motion and (ii) explicit, viewpoint-aware 3D verification prevents semantically plausible but incorrect stops. This suggests that geometry-grounded spatial reasoning is a scalable alternative to heavy policy training or human-in-the-loop interaction for fine-grained instance disambiguation in cluttered 3D scenes.

3D导航空间推理实例定位智能体探索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。