让四足机器人在真实环境里听懂人话自主导航,成功率超88%
Open-Architecture End-to-End System for Real-World Autonomous Robot Navigation
- 用语言指令驱动,融合视觉与语义构建动态场景图
- 基于大模型实时生成并调整多步路径,成功率达88%以上
- 开源架构支持真实部署,适合移动机器人研发者参考
让机器人在未知、复杂且动态的真实环境中自主导航面临诸多挑战,包括感知不完整、观测不全、定位不确定性和安全约束。现有方法多局限于仿真环境,无法应对这些现实问题。本文提出一个轻量级、开放架构的端到端真实世界机器人导航系统。通过在Unitree Go2四足机器人上集成多个机载组件,并利用ROS2进行通信,系统接收自然语言导航指令后,融合机载传感数据进行定位与建图,并结合开放词汇语义构建层级化场景图。基于大模型的规划器利用该图实时生成并适应多步路径规划,随环境变化动态调整。在多个室内环境中测试,实现零样本真实世界自主导航,任务成功率超过88%,并分析了系统实际运行行为。
原文摘要 · Abstract (English)
Enabling robots to autonomously navigate unknown, complex, and dynamic real-world environments presents several challenges, including imperfect perception, partial observability, localization uncertainty, and safety constraints. Current approaches are typically limited to simulations, where such challenges are not present. In this work, we present a lightweight, open-architecture, end-to-end system for real-world robot autonomous navigation. Specifically, we deploy a real-time navigation system on a quadruped robot by integrating multiple onboard components that communicate via ROS2. Given navigation tasks specified in natural language, the system fuses onboard sensory data for localization and mapping with open-vocabulary semantics to build hierarchical scene graphs from a continuously updated semantic object map. An LLM-based planner leverages these graphs to generate and adapt multi-step plans in real time as the scene evolves. Through experiments across multiple indoor environments using a Unitree Go2 quadruped, we demonstrate zero-shot real-world autonomous navigation, achieving over 88% task success, and provide analysis of system behavior during deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。