调整观测空间设计可提升航天器自主控制的强化学习效果
Investigating the Impact of Observation Space Design Choices On Training Reinforcement Learning Solutions for Spacecraft Problems
- 测试不同传感器对智能体学习的影响
- 多数传感器有助于智能体学到更优策略
- 参考坐标系需保持一致,影响较小但不可忽视
近期研究利用强化学习(RL)实现航天器自主控制取得显著进展。然而有研究发现,通过改变动作空间可进一步提升性能,这促使我们探索环境设计的更多优化空间。本文聚焦观测空间设计对航天器巡检任务中强化学习智能体训练与性能的影响。研究分为两部分:一是评估为辅助学习而设计的传感器的作用;二是分析不同参考坐标系对智能体感知世界视角的影响。结果表明,传感器并非必需,但多数情况下有助于智能体学习更优行为;参考坐标系的影响较小,但保持一致更为理想。
原文摘要 · Abstract (English)
Recent research using Reinforcement Learning (RL) to learn autonomous control for spacecraft operations has shown great success. However, a recent study showed their performance could be improved by changing the action space, i.e. control outputs, used in the learning environment. This has opened the door for finding more improvements through further changes to the environment. The work in this paper focuses on how changes to the environment's observation space can impact the training and performance of RL agents learning the spacecraft inspection task. The studies are split into two groups. The first looks at the impact of sensors that were designed to help agents learn the task. The second looks at the impact of reference frames, reorienting the agent to see the world from a different perspective. The results show the sensors are not necessary, but most of them help agents learn more optimal behavior, and that the reference frame does not have a large impact, but is best kept consistent.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。