arXiv:2602.06088cs.LGcs.AI2026-02

用Transformer增强太空碰撞规避,应对观测不全的难题

Transformer-Based Reinforcement Learning for Autonomous Orbital Collision Avoidance in Partially Observable Environments

  • 采用Transformer构建可处理部分可观测性的强化学习框架
  • 在模拟环境中实现高精度碰撞规避,提升复杂观测下的决策可靠性
  • 适合研究空间自主系统与不确定环境交互的学者

我们提出一种基于Transformer的强化学习框架,用于在部分可观测环境下实现自主轨道碰撞规避。该框架结合可配置的交会模拟器、距离相关的观测模型和序列状态估计器,显式建模相对运动中的不确定性。核心贡献是采用基于Transformer的局部可观测马尔可夫决策过程(POMDP)架构,利用长程时序注意力机制,比传统结构更有效地解析噪声大且间歇性的观测数据。该集成方法为训练能在监测不完善环境下可靠运行的规避智能体提供了基础。

原文摘要 · Abstract (English)

We introduce a Transformer-based Reinforcement Learning framework for autonomous orbital collision avoidance that explicitly models the effects of partial observability and imperfect monitoring in space operations. The framework combines a configurable encounter simulator, a distance-dependent observation model, and a sequential state estimator to represent uncertainty in relative motion. A central contribution of this work is the use of transformer-based Partially Observable Markov Decision Process (POMDP) architecture, which leverage long-range temporal attention to interpret noisy and intermittent observations more effectively than traditional architectures. This integration provides a foundation for training collision avoidance agents that can operate more reliably under imperfect monitoring environments.

强化学习轨道避障Transformer不确定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。