用大模型解析文字线索,让机器人更聪明地找房间。
Follow the Signs: Using Textual Cues and LLMs to Guide Efficient Robot Navigation
- 结合视觉感知与大模型,从局部线索推断全局布局
- 在部分可见环境中导航成功率提升25%以上
- 适合需要识字导航的智能机器人场景
在陌生环境中自主导航通常依赖几何地图与规划策略,忽视了标志牌、房间号等丰富的语义线索。本文提出一种新型语义导航框架,利用大语言模型(LLMs)从局部观测中推断模式,预测目标区域最可能的位置。该方法融合局部感知输入与基于前沿的探索,定期查询大模型以提取符号模式(如房间编号规律和建筑布局结构),并更新置信度网格来引导探索。这使得机器人能在未直接观察到目标前,即能高效朝带有文字标识(如“room 8”)的目标位置移动。实验表明,在模拟真实楼层平面的稀疏、部分可观测网格环境中,该方法显著提升导航效率,路径成功率加权路径长度指标优于基线超过25%。
原文摘要 · Abstract (English)
Autonomous navigation in unfamiliar environments often relies on geometric mapping and planning strategies that overlook rich semantic cues such as signs, room numbers, and textual labels. We propose a novel semantic navigation framework that leverages large language models (LLMs) to infer patterns from partial observations and predict regions where the goal is most likely located. Our method combines local perceptual inputs with frontier-based exploration and periodic LLM queries, which extract symbolic patterns (e.g., room numbering schemes and building layout structures) and update a confidence grid used to guide exploration. This enables robots to move efficiently toward goal locations labeled with textual identifiers (e.g., "room 8") even before direct observation. We demonstrate that this approach enables more efficient navigation in sparse, partially observable grid environments by exploiting symbolic patterns. Experiments across environments modeled after real floor plans show that our approach consistently achieves near-optimal paths and outperforms baselines by over 25% in Success weighted by Path Length.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。