arXiv:2507.14731cs.RO2025-07被引 13

一个策略通吃多种机器人,实现跨形态自主导航。

X-Nav: Learning End-to-End Cross-Embodiment Navigation for Mobile Robots

  • 用强化学习训练多个专家策略,再用Transformer蒸馏成通用策略。
  • 在模拟中零样本迁移至未见过的机器人形态和真实感环境。
  • 适合需要快速适配新机器人的移动机器人研发人员。

现有导航方法主要针对特定机器人形态设计,限制了其在不同平台间的泛化能力。本文提出X-Nav,一种端到端跨形态导航框架,使单一统一策略可部署于轮式与四足机器人等多种形态。X-Nav包含两个阶段:1)在大量随机生成的机器人形态上,利用带特权观测的深度强化学习训练多个专家策略;2)通过基于Transformer的导航动作分块(Nav-ACT)从专家策略中蒸馏出单一通用策略。该通用策略直接将视觉与本体感知输入映射为低层控制命令,实现对新形态机器人的泛化。模拟实验表明,X-Nav可在未见形态及逼真环境中实现零样本迁移。可扩展性研究表明,随着训练形态数量增加,性能持续提升。消融实验证明了X-Nav设计的有效性。此外,真实世界实验验证了其在现实环境中的泛化能力。

原文摘要 · Abstract (English)

Existing navigation methods are primarily designed for specific robot embodiments, limiting their generalizability across diverse robot platforms. In this paper, we introduce X-Nav, a novel framework for end-to-end cross-embodiment navigation where a single unified policy can be deployed across various embodiments for both wheeled and quadrupedal robots. X-Nav consists of two learning stages: 1) multiple expert policies are trained using deep reinforcement learning with privileged observations on a wide range of randomly generated robot embodiments; and 2) a single general policy is distilled from the expert policies via navigation action chunking with transformer (Nav-ACT). The general policy directly maps visual and proprioceptive observations to low-level control commands, enabling generalization to novel robot embodiments. Simulated experiments demonstrated that X-Nav achieved zero-shot transfer to both unseen embodiments and photorealistic environments. A scalability study showed that the performance of X-Nav improves when trained with an increasing number of randomly generated embodiments. An ablation study confirmed the design choices of X-Nav. Furthermore, real-world experiments were conducted to validate the generalizability of X-Nav in real-world environments.

机器人导航跨形态强化学习通用策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。