用地球农田训练的智能体,零样本迁移至月球环境导航成功率达近50%。
Transferable Deep Reinforcement Learning for Cross-Domain Navigation: from Farmland to the Moon
- 在农田环境中用PPO训练导航策略,直接迁移到月球模拟场景
- 零样本迁移下月球任务成功率接近50%,无需额外训练
- 为行星探索提供可复用、低成本的自主导航新路径
非结构化环境中的自主导航对野外与行星机器人至关重要,要求在不确定条件下高效抵达目标并避障。传统算法常需针对具体环境大量调参,难以扩展。深度强化学习(DRL)通过与环境直接交互,提供数据驱动的替代方案。本研究探讨在视觉与地形差异显著的模拟域间,基于DRL策略的泛化可行性:在陆地农田环境中训练3D农业探测车的导航策略,采用近端策略优化(PPO)方法实现目标导向导航与障碍规避,并在零样本条件下评估其在类月球环境中的迁移表现。结果表明,地面训练的策略在月球模拟中仍保持较高有效性,成功率接近50%,且无需再训练或微调。这证明跨域DRL策略迁移具备潜力,可为未来行星探测任务提供灵活高效的自主导航方案,同时显著降低重训练成本。
原文摘要 · Abstract (English)
Autonomous navigation in unstructured environments is essential for field and planetary robotics, where robots must efficiently reach goals while avoiding obstacles under uncertain conditions. Conventional algorithmic approaches often require extensive environment-specific tuning, limiting scalability to new domains. Deep Reinforcement Learning (DRL) provides a data-driven alternative, allowing robots to acquire navigation strategies through direct interactions with their environment. This work investigates the feasibility of DRL policy generalization across visually and topographically distinct simulated domains, where policies are trained in terrestrial settings and validated in a zero-shot manner in extraterrestrial environments. A 3D simulation of an agricultural rover is developed and trained using Proximal Policy Optimization (PPO) to achieve goal-directed navigation and obstacle avoidance in farmland settings. The learned policy is then evaluated in a lunar-like simulated environment to assess transfer performance. The results indicate that policies trained under terrestrial conditions retain a high level of effectiveness, achieving close to 50\% success in lunar simulations without the need for additional training and fine-tuning. This underscores the potential of cross-domain DRL-based policy transfer as a promising approach to developing adaptable and efficient autonomous navigation for future planetary exploration missions, with the added benefit of minimizing retraining costs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。