arXiv:2510.20818cs.ROcs.AI2025-10被引 13

VAMOS让机器人在不同地形中导航更可靠,能自动避开无法通过的路径。

VAMOS: A Hierarchical Vision-Language-Action Model for Capability-Modulated and Steerable Navigation

  • 分层设计:高层规划器负责语义路径生成,专用模型评估物理可行性。
  • 真实场景测试中成功率提升3倍,优于现有端到端和基于模型的方法。
  • 支持语言指令控制,可跨轮式与足式机器人通用,适合多形态智能体。

机器人导航的核心挑战在于学习能在多样环境中泛化,同时符合特定本体(如四足机器人可上台阶,而轮式机器人不行)的物理限制。我们提出VAMOS,一种分层视觉-语言-动作模型,将语义规划与本体感知解耦:通用规划器从开放世界数据中学习,专用可行动作模型则在安全、低成本的仿真中学习机器人的物理约束与能力。通过精心设计的接口,高层规划器直接在图像空间提出候选路径,由动作模型评估并重新排序。真实世界实验表明,VAMOS在室内和复杂户外导航中均取得比当前最优基于模型和端到端学习方法更高的成功率。此外,其分层结构支持跨本体导航,可在足式与轮式机器人间通用,并可通过自然语言轻松调控。真实世界消融实验确认,专用模型对本体感知至关重要,使单一高层规划器可部署于物理差异显著的轮式与足式机器人。最终,该模型显著提升单机器人可靠性,通过拒绝物理不可行计划,实现成功率提升3倍。

原文摘要 · Abstract (English)

A fundamental challenge in robot navigation lies in learning policies that generalize across diverse environments while conforming to the unique physical constraints and capabilities of a specific embodiment (e.g., quadrupeds can walk up stairs, but rovers cannot). We propose VAMOS, a hierarchical VLA that decouples semantic planning from embodiment grounding: a generalist planner learns from diverse, open-world data, while a specialist affordance model learns the robot's physical constraints and capabilities in safe, low-cost simulation. We enabled this separation by carefully designing an interface that lets a high-level planner propose candidate paths directly in image space that the affordance model then evaluates and re-ranks. Our real-world experiments show that VAMOS achieves higher success rates in both indoor and complex outdoor navigation than state-of-the-art model-based and end-to-end learning methods. We also show that our hierarchical design enables cross-embodied navigation across legged and wheeled robots and is easily steerable using natural language. Real-world ablations confirm that the specialist model is key to embodiment grounding, enabling a single high-level planner to be deployed across physically distinct wheeled and legged robots. Finally, this model significantly enhances single-robot reliability, achieving 3X higher success rates by rejecting physically infeasible plans. Website: https://vamos-vla.github.io/

机器人导航分层模型多模态跨本体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。