arXiv:2605.15559cs.RO2026-05

提升强化学习机器人导航的仿真到现实迁移效果

NavRL++: A System-Level Framework for Improving Sim-to-Real Transfer in Reinforcement Learning-Based Robot Navigation

论文配图:NavRL++: A System-Level Framework for Improving Sim-to-Real Transfer in Reinforcement Learning-Based Robot Navigation
图 1 · 摘自论文原文
  • 分析仿真到现实迁移的关键影响因素,提出扰动感知微调策略
  • 在静态与动态环境中均优于基线方法,接近优化规划器性能
  • 适用于多类型机器人平台,支持零样本真实部署

近年来,基于强化学习的自主导航取得显著进展。然而,现有方法多关注强化学习框架设计(如输入表示、动作空间、奖励函数),对仿真到现实迁移的分析不足,且缺乏训练策略如何影响真实部署性能的深入理解。为弥补这一差距,我们不仅提出有效的强化学习框架,还构建了完整的训练与部署流水线,并开展系统性实证研究,分离出影响仿真到现实迁移的关键因素:传感器噪声、感知失败、系统延迟和控制响应。基于分析结果,我们引入扰动感知微调,一种显式考虑实测域差异的后训练适配策略,以增强迁移鲁棒性。为进一步缓解真实场景中的感知退化并提升控制平滑性,我们提出基于Transformer的时序推理策略,利用短时观察进行导航控制。我们在多个环境上定量评估各仿真到现实扰动及训练设计选择对导航性能的影响。实验表明,所提训练策略与模型架构在静态与动态环境中均优于基于学习的基线方法,且在静态场景中性能可媲美优化类规划器。通过在多种机器人平台(包括飞行与腿式机器人)上的真实部署验证,涵盖探索与巡检等导航任务,实现了零样本仿真到现实迁移。

原文摘要 · Abstract (English)

Recent years have witnessed significant progress in autonomous navigation using reinforcement learning. However, existing approaches largely emphasize reinforcement learning framework design, such as input representations, action spaces, and reward functions, while providing limited analysis of sim-to-real transfer and insufficient insight into how training strategies affect real-world deployment performance. To bridge this gap, we not only introduce an effective RL framework but also present a complete training and deployment pipeline, along with a systematic empirical study that disentangles the key factors affecting sim-to-real transfer in reinforcement learning-based navigation, including sensor noise, perception failures, system latency, and control response. Building on insights from this analysis, we introduce perturbation-aware fine-tuning, a post-training adaptation strategy that improves transfer robustness by explicitly accounting for empirically identified domain discrepancies. To further mitigate perception degradation and enhance control smoothness in real-world deployment, we propose a Transformer-based temporal reasoning policy that leverages short-horizon observation for navigation control. We quantitatively evaluate how individual sim-to-real perturbations and training design choices impact navigation performance across environments. Experimental results demonstrate that the proposed training strategy and policy architecture outperform learning-based baselines in both static and dynamic environments, while achieving performance comparable to optimization-based planners in static settings. We validate our approach through real-world deployment on multiple robotic platforms, including aerial and legged robots, across navigation-centric tasks such as exploration and inspection, demonstrating zero-shot sim-to-real transfer.

强化学习机器人导航仿真到现实迁移学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。