提出脑-小脑架构,实现跨机器人平台的快速安全零样本导航。
FSUNav: A Cerebrum-Cerebellum Architecture for Fast, Safe, and Universal Zero-Shot Goal-Oriented Navigation
- 采用脑-小脑双模块设计,小脑负责高速局部规划,脑部处理语义推理。
- 在多个基准上实现领先性能,真实机器人部署验证了鲁棒性。
- 支持多模态输入,适用于开放词汇目标导航,适合通用机器人系统。
当前视觉-语言导航方法在异构机器人兼容性、实时性能和导航安全性方面存在显著瓶颈,且难以支持开放词汇语义泛化与多模态任务输入。为此,本文提出FSUNav:一种用于快速、安全、通用零样本目标导向导航的脑-小脑架构,创新性地将视觉-语言模型(VLMs)与新架构结合。小脑模块作为高频端到端模块,基于深度强化学习构建通用局部规划器,实现异构平台(如人形、四足、轮式机器人)上的统一导航,提升效率并显著降低碰撞风险。脑模块构建三层推理模型,利用VLMs建立端到端检测与验证机制,实现无需预定义标识的零样本开放词汇目标导航,在仿真与真实环境均提升任务成功率。此外,该框架支持多模态输入(如文本、目标描述、图像),进一步增强泛化能力、实时性、安全性与鲁棒性。在MP3D、HM3D和OVON基准上的实验结果表明,FSUNav在物体、实例图像及任务导航上均达到最先进水平,显著优于现有方法。多种机器人平台的真实世界部署进一步验证其鲁棒性与实际应用价值。
原文摘要 · Abstract (English)
Current vision-language navigation methods face substantial bottlenecks regarding heterogeneous robot compatibility, real-time performance, and navigation safety. Furthermore, they struggle to support open-vocabulary semantic generalization and multimodal task inputs. To address these challenges, this paper proposes FSUNav: a Cerebrum-Cerebellum architecture for fast, safe, and universal zero-shot goal-oriented navigation, which innovatively integrates vision-language models (VLMs) with the proposed architecture. The cerebellum module, a high-frequency end-to-end module, develops a universal local planner based on deep reinforcement learning, enabling unified navigation across heterogeneous platforms (e.g., humanoid, quadruped, wheeled robots) to improve navigation efficiency while significantly reducing collision risk. The cerebrum module constructs a three-layer reasoning model and leverages VLMs to build an end-to-end detection and verification mechanism, enabling zero-shot open-vocabulary goal navigation without predefined IDs and improving task success rates in both simulation and real-world environments. Additionally, the framework supports multimodal inputs (e.g., text, target descriptions, and images), further enhancing generalization, real-time performance, safety, and robustness. Experimental results on MP3D, HM3D, and OVON benchmarks demonstrate that FSUNav achieves state-of-the-art performance on object, instance image, and task navigation, significantly outperforming existing methods. Real-world deployments on diverse robotic platforms further validate its robustness and practical applicability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。