用图注意力机制让月球探测机器人自主路由,不依赖全局信息却能高效传数据。
Learning Decentralized Routing Policies via Graph Attention-based Multi-Agent Reinforcement Learning in Lunar Delay-Tolerant Networks
- 基于图注意力的多智能体强化学习,仅靠局部信息决策。
- 在随机探索场景中交付率高、无重复、丢包少,支持短期移动预测。
- 适合未来大规模月球探测机器人集群,可扩展性强。
我们提出一种完全去中心化的路由框架,用于在月球延迟容忍网络(LDTN)约束下进行多机器人探索任务。在此场景中,自主探测车需在间歇性连接和未知移动模式下,将采集数据中继至着陆器。我们将问题建模为部分可观测马尔可夫决策过程(POMDP),并提出基于图注意力的多智能体强化学习(GAT-MARL)策略,采用集中训练、分散执行(CTDE)范式。该方法仅依赖局部观测,无需全局拓扑更新或包复制,区别于传统最短路径和受控洪泛算法。通过在随机探索环境中的蒙特卡洛仿真,GAT-MARL展现出更高的数据交付率、零冗余与更少丢包,并能利用短期移动预测;在更大规模探测车团队上也表现良好,验证了其在行星探测未来空间机器人系统中的可扩展性。
原文摘要 · Abstract (English)
We present a fully decentralized routing framework for multi-robot exploration missions operating under the constraints of a Lunar Delay-Tolerant Network (LDTN). In this setting, autonomous rovers must relay collected data to a lander under intermittent connectivity and unknown mobility patterns. We formulate the problem as a Partially Observable Markov Decision Problem (POMDP) and propose a Graph Attention-based Multi-Agent Reinforcement Learning (GAT-MARL) policy that performs Centralized Training, Decentralized Execution (CTDE). Our method relies only on local observations and does not require global topology updates or packet replication, unlike classical approaches such as shortest path and controlled flooding-based algorithms. Through Monte Carlo simulations in randomized exploration environments, GAT-MARL provides higher delivery rates, no duplications, and fewer packet losses, and is able to leverage short-term mobility forecasts; offering a scalable solution for future space robotic systems for planetary exploration, as demonstrated by successful generalization to larger rover teams.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。