arXiv:2606.10771astro-ph.IMcs.LG2026-06中稿 · A&A

首次在真实天文观测中验证强化学习控制器,显著提升自适应光学性能。

On-sky demonstration of reinforcement learning for adaptive optics control

论文配图:On-sky demonstration of reinforcement learning for adaptive optics control
图 1 · 摘自论文原文
  • 用强化学习设计控制器PO4AO,通过在线学习补偿振动和噪声。
  • 实测表现优于传统积分控制器,跨不同条件保持稳定效果。
  • 适合希望实现高鲁棒性自适应光学系统的天文台或研发团队。

基于强化学习(RL)的算法近年来被视为自适应光学(AO)控制的有前景方案。尽管在模拟和实验室中已展示对光子噪声、探测器噪声、错位、振动及快速变化的大气条件的鲁棒性,但其在真实天空环境中的表现尚未验证。本文报告了首个面向实际天文观测的强化学习控制器——用于自适应光学的策略优化(PO4AO)的在轨演示。该系统部署于位于上普罗旺斯天文台(OHP)1.52米望远镜(T152)卡塞格林焦面的Papyrus AO系统,通过共享内存接口与现有实时控制器(DAO RTC)集成。在多个夜晚的观测中,对不同通量水平和大气条件下的性能进行了对比测试。结果表明,无论何种配置,PO4AO始终优于标准积分控制器。该控制器成功学习并补偿了振动模式,对测量噪声表现出强鲁棒性。一旦针对Papyrus完成调参,即可在不调整参数的情况下,适应多种观测条件和科学目标。这些性能提升是在未优化的Python实现引入约750 μs额外延迟、控制抖动和偶尔帧丢失的条件下达成的。若经合理优化,PO4AO将成为单共轭自适应光学系统中鲁棒且高性能的即插即用控制器,推动强化学习在真实天文观测中的广泛应用。

原文摘要 · Abstract (English)

Reinforcement learning (RL)-based algorithms have recently emerged as a promising approach for adaptive optics (AO) control. In simulations and laboratory experiments, they have demonstrated robustness to real-world effects such as photon and detector noise, misregistration, vibrations, and rapid variations in seeing conditions. However, their performance has not yet been validated on sky. We report the first on-sky demonstration of a reinforcement learning controller for adaptive optics, named Policy Optimization for AO (PO4AO). We further analyze its on-sky behavior and identify directions for improving the algorithm and its implementation.PO4AO was implemented and deployed on the Papyrus adaptive optics system installed at the Coudé focus of the 1.52 m telescope (T152) at the OHP. A Python-based implementation was interfaced with the existing real-time controller (DAO RTC) via shared-memory buffers. The performance of PO4AO was compared to that of a standard integrator controller over several nights, covering a range of flux levels and atmospheric conditions. PO4AO consistently outperformed the standard integrator in all tested configurations. The controller successfully learned and compensated for vibration patterns and demonstrated strong robustness to measurement noise. Once tuned for Papyrus, PO4AO operated in a turnkey fashion, using a single set of hyperparameters across varying observing conditions and science targets. These performance gains were achieved despite a non-optimized Python implementation introducing approximately $750\,μ\text{s}$ of additional latency, along with control jitter and occasional frame drops. When properly implemented and optimized, PO4AO constitutes a robust and high-performance turnkey controller for single-conjugate adaptive optics systems, paving the way for broader adoption of reinforcement learning strategies in on-sky AO operations.

自适应光学强化学习天文观测实时控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。