用强化学习精准调控激光脉冲,实现自适应高精度控制。
Shaping Laser Pulses with Reinforcement Learning
- 仅用非破坏性图像观测,通过强化学习学控制策略。
- 在动态变化环境下仍能达目标强度的90%。
- 适合需要实时适应复杂物理系统的激光实验人员。
高功率激光(HPL)系统工作在阿秒量级——人类迄今创造的最短时间尺度。这些系统在高能物理中至关重要,利用超短脉冲持续时间实现极高强度,对光与物质相互作用的实际应用和理论突破均具关键意义。传统上,调节HPL光学性能依赖人工专家调参,或采用计算成本高的黑箱优化方法。但后者依赖静态假设,忽视高能物理中的复杂动态及日常实验环境的变化,常需频繁重启。深度强化学习(DRL)提供了一种新可能,可在非静态环境中实现序列决策。本文探索了将DRL应用于HPL系统的可行性,拓展现有研究:(1)仅依赖可得诊断设备提供的非破坏性图像观测,学习控制策略;(2)在系统动态变化时保持性能。我们在多种测试动态下评估该方法,发现DRL有效实现跨域适应,在动态波动中仍达到目标强度的90%。
原文摘要 · Abstract (English)
High Power Laser (HPL) systems operate in the attoseconds regime -- the shortest timescale ever created by humanity. HPL systems are instrumental in high-energy physics, leveraging ultra-short impulse durations to yield extremely high intensities, which are essential for both practical applications and theoretical advancements in light-matter interactions. Traditionally, the parameters regulating HPL optical performance have been manually tuned by human experts, or optimized using black-box methods that can be computationally demanding. Critically, black box methods rely on stationarity assumptions overlooking complex dynamics in high-energy physics and day-to-day changes in real-world experimental settings, and thus need to be often restarted. Deep Reinforcement Learning (DRL) offers a promising alternative by enabling sequential decision making in non-static settings. This work explores the feasibility of applying DRL to HPL systems, extending the current research by (1) learning a control policy relying solely on non-destructive image observations obtained from readily available diagnostic devices, and (2) retaining performance when the underlying dynamics vary. We evaluate our method across various test dynamics, and observe that DRL effectively enables cross-domain adaptability, coping with dynamics' fluctuations while achieving 90\% of the target intensity in test environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。