arXiv:2605.19206cs.RO2026-05

根据目标特性动态调整房间与物体线索权重,提升零样本导航效率。

CLUE: Adaptively Prioritized Contextual Cues by Leveraging a Unified Semantic Map for Effective Zero-Shot Object-Goal Navigation

论文配图:CLUE: Adaptively Prioritized Contextual Cues by Leveraging a Unified Semantic Map for Effective Zero-Shot Object-Goal Navigation
图 1 · 摘自论文原文
  • 用大模型提取常识知识,判断目标与房间的关联强度
  • 构建加权语义地图,在模拟与真实环境均优于基线
  • 适合需要快速适应新场景的机器人导航任务

零样本物体目标导航(ZSON)是机器人领域中的挑战性问题,需同时理解语言与视觉信息。房间和物体的上下文线索至关重要,但其重要性取决于目标类型:某些物体与特定房间强相关,而另一些则更依赖邻近物体。现有方法忽视此差异,导致探索低效且不准确。我们提出CLUE框架,通过离线大语言模型(LLM)提取常识知识,估计目标与房间类型的关联度,从而自适应地优先使用房间线索(对可预测物体)或物体线索(对弱房间关联物体)。该框架构建统一的语义值地图,依据目标模糊性动态加权两类线索以指导探索。结合多视角验证与基于上下文线索的探索策略,实验表明在仿真与真实世界部署中,本方法在成功率(SR)与路径长度加权成功率(SPL)上均持续优于现有最优基线,验证了其有效性和实用性。

原文摘要 · Abstract (English)

Zero-shot object-goal navigation (ZSON) is a challenging problem in robotics that requires a comprehensive understanding of both language and visual observations. Contextual cues from rooms and objects are critical, but their relative importance depends on the target: some objects are strongly tied to specific room types, while others are better predicted by nearby co-located objects. Existing methods overlook this distinction, leading to inefficient and inaccurate exploration. We present CLUE, a novel navigation framework that adaptively balances the use of contextual rooms and objects by leveraging commonsense knowledge extracted from an offline large language model (LLM). By estimating a target's association with room types using LLM, the agent prioritizes room cues for predictable objects and object cues for those with weak room associations. Our framework constructs a unified semantic value map that integrates both types of contextual information, adaptively weighted by the target's ambiguity to guide exploration. Combined with multi-viewpoint verification and an exploration strategy informed by contextual cues, CLUE achieves robust and efficient navigation. Extensive experiments in simulation and real-world deployments show that our method consistently outperforms state-of-the-art baselines in both success rate (SR) and success weighted by path length (SPL), demonstrating its effectiveness and practicality for real-world navigation tasks.

机器人导航零样本语义地图大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。