arXiv:2605.05857cs.LG2026-05被引 1

用历史数据训练强化学习,实现托卡马克等离子体旋转剖面精准控制。

Offline Reinforcement Learning for Rotation Profile Control in Tokamaks

论文配图:Offline Reinforcement Learning for Rotation Profile Control in Tokamaks
图 1 · 摘自论文原文
  • 基于离线强化学习与概率动力学模型,仅用历史数据训练控制策略。
  • 在DIII-D托卡马克上部署后,实测结果展现良好控制效果。
  • 适合研究核聚变控制、强化学习应用的科研人员参考。

托卡马克仍是实现可控核聚变能源的主要候选装置,但其内部诸多控制问题仍难以解决。其中,等离子体旋转剖面控制尤为关键,因其显著影响稳定性、约束性能和输运特性。尽管平均旋转可被调控,但完整剖面控制因高维性、多执行器响应及对等离子体状态的依赖而极具挑战。基于学习的控制方法(如强化学习)具备建模复杂交互的能力,可实现多输入多输出的有效控制。然而,由于缺乏能准确模拟旋转剖面动态的仿真器,学习此类策略十分困难。本文研究了离线强化学习及离线模型基础强化学习算法在该任务中的应用,仅使用DIII-D托卡马克的历史数据进行训练。最终方法采用等离子体动力学的概率模型生成轨迹用于强化学习训练。将该策略部署于DIII-D托卡马克并获得有前景的现实实验结果。最后,我们总结了在仅依赖有限历史数据的情况下,在复杂物理装置上训练与部署强化学习策略所面临的挑战与关键洞见。

原文摘要 · Abstract (English)

Tokamaks remain leading candidates for achieving practical fusion energy, yet many important control problems inside these devices are still difficult or unsolved. One such challenge is controlling the plasma rotation profile, which strongly influences stability, confinement, and transport. While the average rotation can be controlled, controlling the full profile is challenging due to high dimensionality, response to multiple actuators and dependence on plasma condition. Learning-based control methods, such as reinforcement learning (RL), provide a potential solution to this challenging problem with ability to model complex interactions leading to effective multi-input multi-output control. However, learning such policies is challenging due to the lack of accurate simulators that can model the rotation profile dynamics. In this work, we investigate the use of offline RL and offline model-based RL algorithms for rotation profile control, training them solely on historical data from the DIII-D tokamak. Our final method uses probabilistic models of plasma dynamics to generate rollouts for RL training. We deploy this policy on the DIII-D Tokamak and observe promising real-world results. We conclude by highlighting key challenges and insights from training and deploying an RL policy on a complex physical device while using only limited past data.

强化学习核聚变控制理论离线学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。