arXiv:2506.07338cs.CVcs.RO2025-06被引 3

用分层评分选最优视角,让机器人更高效找目标物。

Hierarchical Scoring with 3D Gaussian Splatting for Instance Image-Goal Navigation

  • 分层评分:先用语义相似度筛选区域,再精细估算姿态。
  • 在模拟和真实场景中均达到顶尖性能,减少冗余渲染。
  • 适合需要精准导航的机器人应用,如智能客服或仓储巡检。

实例图像目标导航(IIN)要求自主代理识别并导航至参考图像中任意视角拍摄的目标物体或位置。尽管近期方法利用强大的新视图合成技术(如3D高斯点云,3DGS),但通常依赖随机采样多个视角或路径以确保关键视觉线索的全面覆盖。这种方法导致图像样本重叠严重,缺乏有原则的视角选择,显著增加渲染与比对开销。本文提出一种新型IIN框架,采用分层评分机制,估计目标匹配的最佳视角。该方法结合跨层级语义评分(利用CLIP生成的相关性场识别与目标类别高度相似的区域)与细粒度局部几何评分(在有前景区域内进行精确位姿估计)。大量实验表明,本方法在模拟IIN基准上表现卓越,并具备实际应用场景可行性。

原文摘要 · Abstract (English)

Instance Image-Goal Navigation (IIN) requires autonomous agents to identify and navigate to a target object or location depicted in a reference image captured from any viewpoint. While recent methods leverage powerful novel view synthesis (NVS) techniques, such as three-dimensional Gaussian splatting (3DGS), they typically rely on randomly sampling multiple viewpoints or trajectories to ensure comprehensive coverage of discriminative visual cues. This approach, however, creates significant redundancy through overlapping image samples and lacks principled view selection, substantially increasing both rendering and comparison overhead. In this paper, we introduce a novel IIN framework with a hierarchical scoring paradigm that estimates optimal viewpoints for target matching. Our approach integrates cross-level semantic scoring, utilizing CLIP-derived relevancy fields to identify regions with high semantic similarity to the target object class, with fine-grained local geometric scoring that performs precise pose estimation within promising regions. Extensive evaluations demonstrate that our method achieves state-of-the-art performance on simulated IIN benchmarks and real-world applicability.

机器人导航3D高斯视觉定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。