用强化学习精准控制核聚变装置中X点位置,提升稳定性和鲁棒性。
Advantage-level Aggregation Reinforcement Learning for X-point Target Magnetic Configuration Control in an EXL-50U Experiment-Calibrated Simulation Environment

- 提出优势聚合方法,解决多目标控制中奖励混淆问题。
- 在500毫秒内将最差通道得分从0.23提升至0.81,误差降低20倍。
- 可在未知初始状态和测量噪声下稳定运行,适合真实装置部署。
高功率紧凑型托卡马克面临分流器热负载管理难题。为增强局部通量扩展并分离耗散区与核心区域,EHL-2采用X点靶(XPT)分流器。这要求次级X点始终位于分流腿上;偏移会破坏拓扑结构与排气几何形状。当前实验(包括EXL-50U放电)依赖预计算前馈波形配合全局量的PID环路,缺乏对次级零点的专用闭环反馈,导致XPT操作可重复但难常态化。本文在与EXL-50U放电#13906校准的自由边界环境中,将XPT反馈建模为多目标强化学习(RL)控制问题。针对等离子体电流、形状与零点约束间的强耦合——传统奖励标量化导致目标时间信用丢失——我们提出优势聚合(AdvA)。AdvA在最差目标感知的非线性标量化前保留各目标的时间信用,并引入残差修正项优化策略更新。使用AdvA-PPO在正常工况、测量不确定性及未见初始平衡条件下,对比Reward-PPO与前馈+PID基线进行评估。在500毫秒滚动周期内,AdvA-PPO将平均最差通道得分从0.23提升至0.81,降低X点通量均方根误差约20倍。在综合测量不确定性下,它是唯一能完成整个时域且维持可用XPT形态的学习控制器。多初始状态微调使单一AdvA-PPO策略可在分流器与限流器初始平衡状态下实现全时域运行。这些结果为未来在EXL-50U上开展实时XPT验证提供了仿真基础。
原文摘要 · Abstract (English)
Managing divertor heat loads is a central challenge for compact, high-power tokamaks. To increase local flux expansion and decouple the dissipation volume from the core, EHL-2 adopts the X-point target (XPT) divertor. This requires the secondary X-point to remain on the divertor leg; displacement degrades the topology and exhaust geometry. Current experiments, including EXL-50U discharges, rely on precomputed feedforward waveforms with PID loops on global quantities. Lacking dedicated closed-loop feedback for the secondary null, XPT operation is repeatable but not routine. We formulate XPT feedback as a multi-objective reinforcement learning (RL) control problem in a free-boundary environment calibrated to EXL-50U discharge #13906. To address strong coupling among plasma current, shape, and null constraints - where reward scalarisation collapses objective-specific temporal credit - we develop Advantage Aggregation (AdvA). AdvA preserves objective-wise temporal credit before worst-objective-aware nonlinear scalarisation and introduces a residual correction to policy updates. AdvA-PPO is evaluated against Reward-PPO and a feedforward-plus-PID baseline under nominal operation, measurement uncertainties, and unseen initial equilibria. On a 500 ms rollout, AdvA-PPO raises the mean worst-channel score from 0.23 to 0.81 over Reward-PPO, reducing X-point flux RMSE by ~20x. Under combined measurement uncertainties, it is the only learned controller completing the horizon while retaining a usable XPT shape. Multi-initialization fine-tuning enables a single AdvA-PPO policy to complete full-horizon operation across divertor and limiter initial equilibria. These results provide a simulation-based foundation for future real-time XPT validation on EXL-50U.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。