arXiv:2409.17443cs.RO2024-09被引 5

用对抗强化学习训练卫星躲避多敌方追击,提升太空自主应变能力。

Satellite Chasers: Divergent Adversarial Reinforcement Learning to Engage Intelligent Adversaries on Orbit

  • 分两阶段训练,通过多样化敌方策略增强探索,提升规避模型鲁棒性。
  • 在猫鼠卫星对战中表现优于传统路径规划方法,有效应对多敌围追。
  • 适合研究太空自主博弈、智能对抗系统的科研人员和工程师参考。

随着太空日益拥挤且充满竞争,多智能体环境下的自主能力变得至关重要。当前空间自主系统主要依赖基于优化的路径规划或远距离轨道机动,在一方主动追击另一方的对抗场景中尚未证明有效。本文提出分阶段多智能体强化学习方法DARL(Divergent Adversarial Reinforcement Learning),用于训练卫星在多个敌方航天器追击下的自主规避策略。该方法通过促进多样化对抗策略来增强训练过程中的探索能力,从而生成更鲁棒、适应性更强的避让模型。我们在一个模拟‘猫鼠’卫星对抗的不完全可观测多智能体捕旗游戏中验证了DARL,其中两艘敌方‘猫’航天器追逐一艘‘鼠’逃逸者。实验对比了包括基于优化的卫星路径规划在内的多个基准方法,结果表明DARL能有效生成适用于对抗性多智能体太空环境的高鲁棒性模型。

原文摘要 · Abstract (English)

As space becomes increasingly crowded and contested, robust autonomous capabilities for multi-agent environments are gaining critical importance. Current autonomous systems in space primarily rely on optimization-based path planning or long-range orbital maneuvers, which have not yet proven effective in adversarial scenarios where one satellite is actively pursuing another. We introduce Divergent Adversarial Reinforcement Learning (DARL), a two-stage Multi-Agent Reinforcement Learning (MARL) approach designed to train autonomous evasion strategies for satellites engaged with multiple adversarial spacecraft. Our method enhances exploration during training by promoting diverse adversarial strategies, leading to more robust and adaptable evader models. We validate DARL through a cat-and-mouse satellite scenario, modeled as a partially observable multi-agent capture the flag game where two adversarial ``cat" spacecraft pursue a single ``mouse" evader. DARL's performance is compared against several benchmarks, including an optimization-based satellite path planner, demonstrating its ability to produce highly robust models for adversarial multi-agent space environments.

强化学习太空博弈多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。