为核聚变等离子体控制构建首个离线强化学习基准,支持多任务闭环评估。
Offline Reinforcement Learning for Plasma Control in Nuclear Fusion: Codebase and Benchmark

- 基于DIII-D实测数据构建融合等离子体控制环境,支持四类轨迹跟踪任务。
- 模型驱动的离线强化学习方法在多数任务上表现最佳,凸显动态建模重要性。
- 开源代码、数据与框架,推动聚变控制与离线强化学习共同发展。
离线强化学习(Offline RL)为从历史托卡马克数据中开发等离子体控制器提供了前景,因真实设备上的在线试错成本高且风险大。然而,由于缺乏针对核聚变中复杂多执行器、长时序控制问题的标准离线RL基准,相关进展难以衡量。我们提出RL4F:面向核聚变等离子体控制的离线强化学习基准,提供闭环评估环境和四类完整剖面跟踪任务(旋转、密度、温度、压力)的基线对比。评估环境的动力学模型基于真实DIII-D托卡马克的历史放电数据构建。我们在统一协议下评估多种模仿学习与离线强化学习基线方法。结果表明,模型驱动的离线强化学习在多数目标上取得最优平均性能,但无单一方法在所有任务上全面领先,凸显复杂长时序控制中动态建模的关键作用。为促进后续研究,我们开源代码库、数据集与评估框架,为聚变领域及离线强化学习算法发展提供通用基准。
原文摘要 · Abstract (English)
Offline reinforcement learning (RL) offers a promising route for developing plasma controllers from historical tokamak data, since online trial-and-error on real devices is costly and risky. However, progress in this direction remains difficult to measure due to the lack of a standardized offline RL benchmark for realistic multi-actuator, long-horizon plasma control problems in nuclear fusion. We introduce RL4F, an Offline Reinforcement Learning Benchmark for Plasma Control in Nuclear Fusion, providing closed-loop evaluation environments and baseline comparisons across four full-profile tracking tasks: rotation, density, temperature, and pressure. The dynamics function underlying the evaluation environment is built from historical discharge data from DIII-D, a real-world Tokamak. We evaluate a broad set of imitation learning and offline RL baselines under a unified protocol. We find that offline model-based RL methods obtain the best average performance on most objectives, although no single method dominates all tasks, highlighting the importance of dynamics modeling in complex, long-horizon plasma control tasks. To foster further research, we open-source the codebase, datasets, and evaluation framework, providing a benchmark not only for the fusion community but also for algorithm development in offline RL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。