让机器人在家中自动发现物品并适应环境变化,搜索效率提升超29%。
Where Did I Leave My Glasses? Open-Vocabulary Semantic Exploration in Real-World Semi-Static Environments
- 用概率模型跟踪物体是否静止,主动探测长时间未访问区域。
- 结合大模型理解语义,搜索时优先考虑相关区域,成功率更高。
- 真实场景验证有效,平均检测95%的环境变化,适合家庭服务机器人。
部署于家庭等真实环境中的机器人不仅需安全导航,还需理解周围世界并适应变化。为高效完成任务,必须构建并维护反映当前环境状态的语义地图。现有研究多聚焦静态场景,缺乏对物体实例的持续追踪。本文提出一种面向半静态环境的开放词汇语义探索系统,通过建立物体实例稳定性的概率模型,系统性追踪半静态变化,并主动探索长期未访问区域,实现一致地图维护。同时,利用大语言模型(LLM)进行语义推理,支持开放词汇的目标物体导航,使机器人能基于上下文优先搜索相关区域,提升搜索效率。我们在公开的物体导航与建图数据集上对比了先进基线方法,并在三个真实环境中验证了迁移能力。结果表明,本方法在任务成功率和搜索效率上均优于基线,在处理半静态环境变化时更具鲁棒性;真实实验中平均检测到95%的环境变化,效率比随机和巡检策略提升超过29%。
原文摘要 · Abstract (English)
Robots deployed in real-world environments, such as homes, must not only navigate safely but also understand their surroundings and adapt to changes in the environment. To perform tasks efficiently, they must build and maintain a semantic map that accurately reflects the current state of the environment. Existing research on semantic exploration largely focuses on static scenes without persistent object-level instance tracking. In this work, we propose an open-vocabulary, semantic exploration system for semi-static environments. Our system maintains a consistent map by building a probabilistic model of object instance stationarity, systematically tracking semi-static changes, and actively exploring areas that have not been visited for an extended period. In addition to active map maintenance, our approach leverages the map's semantic richness with large language model (LLM)-based reasoning for open-vocabulary object-goal navigation. This enables the robot to search more efficiently by prioritizing contextually relevant areas. We compare our approach against state-of-the-art baselines using publicly available object navigation and mapping datasets, and we further demonstrate real-world transferability in three real-world environments. Our approach outperforms the compared baselines in both success rate and search efficiency for object-navigation tasks and can more reliably handle changes in mapping semi-static environments. In real-world experiments, our system detects 95% of map changes on average, improving efficiency by more than 29% as compared to random and patrol strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。