arXiv:2411.12183eess.SYcs.LG2024-11被引 1

用强化学习自动对准光束线,提升精度与效率。

Action-Attentive Deep Reinforcement Learning for Autonomous Alignment of Beamlines

  • 将对准问题建模为马尔可夫决策过程,用带动作注意力的策略网络优化调整决策。
  • 在两个模拟光束线上实验,性能优于贝叶斯优化和传统强化学习方法。
  • 适合需要高精度光束控制的同步辐射实验领域研究者参考。

同步辐射光源在材料科学、生物学和化学等领域至关重要。光束线作为其关键子系统,负责调控并引导辐射至样品进行分析。然而,光束线对准过程复杂且耗时,主要依赖经验工程师手动完成。即使光学元件存在微小偏移,也会显著影响光束特性,导致实验结果不理想。现有自动化方法如贝叶斯优化(BO)和强化学习(RL)虽有所改进,但仍存在局限:未充分考虑当前与目标光束状态之间的关系,且忽视光学元件的物理特性,例如需调节特定器件以控制输出光束的斑点尺寸或位置。本文将光束线对准建模为马尔可夫决策过程(MDP),通过强化学习训练智能体。该智能体根据当前与目标光束状态计算调整值,执行动作并迭代优化直至达到最优参数。设计了带动作注意力的策略网络,综合考虑状态差异与光学元件影响,提升决策能力。在两个模拟光束线上的实验表明,所提算法优于现有方法;消融实验进一步验证了动作注意力机制的有效性。

原文摘要 · Abstract (English)

Synchrotron radiation sources play a crucial role in fields such as materials science, biology, and chemistry. The beamline, a key subsystem of the synchrotron, modulates and directs the radiation to the sample for analysis. However, the alignment of beamlines is a complex and time-consuming process, primarily carried out manually by experienced engineers. Even minor misalignments in optical components can significantly affect the beam's properties, leading to suboptimal experimental outcomes. Current automated methods, such as bayesian optimization (BO) and reinforcement learning (RL), although these methods enhance performance, limitations remain. The relationship between the current and target beam properties, crucial for determining the adjustment, is not fully considered. Additionally, the physical characteristics of optical elements are overlooked, such as the need to adjust specific devices to control the output beam's spot size or position. This paper addresses the alignment of beamlines by modeling it as a Markov Decision Process (MDP) and training an intelligent agent using RL. The agent calculates adjustment values based on the current and target beam states, executes actions, and iterates until optimal parameters are achieved. A policy network with action attention is designed to improve decision-making by considering both state differences and the impact of optical components. Experiments on two simulated beamlines demonstrate that our algorithm outperforms existing methods, with ablation studies highlighting the effectiveness of the action attention-based policy network.

强化学习光束对准同步辐射智能控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。