arXiv:2502.05069cs.RO2025-02被引 7

用深度强化学习提升地磁导航跨区域适应能力

Exploring the Generalizability of Geomagnetic Navigation: A Deep Reinforcement Learning approach with Policy Distillation

  • 多教师策略蒸馏融合不同区域的地磁导航策略
  • 跨域导航成功率提升,路径更短、偏差更小
  • 适合需要无GPS环境下长期稳定导航的系统

自动驾驶车辆在未知环境中的导航与探索能力不断提升。地磁导航因其不依赖GPS或惯性设备而受到关注。尽管相关方法已广泛研究,但学习到的地磁导航策略在源域外的表现仍缺乏探索。由于对新进入区域地磁特征缺乏了解,策略性能会下降。本文通过深度强化学习(DRL)探索地磁导航策略的泛化能力:从多个分布区域训练多个教师模型,代表分散的导航策略;设计结合基于势能和内在动机的奖励机制,提升探索效率与模型表征能力;再通过多教师策略蒸馏整合各教师策略,形成跨区域通用的导航方案。数值仿真结果表明,所提方法可有效将DRL模型迁移到新区域,相比现有进化类方法,在导航长度、时长、航向偏差和成功率方面均表现更优。

原文摘要 · Abstract (English)

The advancement in autonomous vehicles has empowered navigation and exploration in unknown environments. Geomagnetic navigation for autonomous vehicles has drawn increasing attention with its independence from GPS or inertial navigation devices. While geomagnetic navigation approaches have been extensively investigated, the generalizability of learned geomagnetic navigation strategies remains unexplored. The performance of a learned strategy can degrade outside of its source domain where the strategy is learned, due to a lack of knowledge about the geomagnetic characteristics in newly entered areas. This paper explores the generalization of learned geomagnetic navigation strategies via deep reinforcement learning (DRL). Particularly, we employ DRL agents to learn multiple teacher models from distributed domains that represent dispersed navigation strategies, and amalgamate the teacher models for generalizability across navigation areas. We design a reward shaping mechanism in training teacher models where we integrate both potential-based and intrinsic-motivated rewards. The designed reward shaping can enhance the exploration efficiency of the DRL agent and improve the representation of the teacher models. Upon the gained teacher models, we employ multi-teacher policy distillation to merge the policies learned by individual teachers, leading to a navigation strategy with generalizability across navigation domains. We conduct numerical simulations, and the results demonstrate an effective transfer of the learned DRL model from a source domain to new navigation areas. Compared to existing evolutionary-based geomagnetic navigation methods, our approach provides superior performance in terms of navigation length, duration, heading deviation, and success rate in cross-domain navigation.

地磁导航强化学习策略蒸馏跨域泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。