arXiv:2605.15935cs.ROcs.SY2026-05

用强化学习实现磁约束聚变中动态形状的实时控制,支持传感器部分失效。

Dynamic Plasma Shape Control with Arbitrary Sensor Subsets

论文配图:Dynamic Plasma Shape Control with Arbitrary Sensor Subsets
图 1 · 摘自论文原文
  • 基于强化学习的统一控制器,直接从原始传感器数据输出线圈指令。
  • 在模拟中动态跟踪误差仅2.01厘米,30%传感器失效下仍稳定运行。
  • 可零样本迁移至真实装置和不同仿真环境,适合聚变控制研究者。

托卡马克中的等离子体形状控制需要实时控制器在动态变化的目标下运行,并容忍诊断设备故障。传统方法将问题分解为平衡重建与线性控制两步,且假设传感器组始终完整可用。本文提出一种强化学习代理,同时解决上述两个限制。该代理在针对DIII-D配置的高保真度NSFsim模拟器中训练,使用包含120种实验等离子体形状的精选数据集。形状目标每0.25秒随机跳变,使代理暴露于全形状空间内的多样化过渡。测试时,代理零样本即可追踪动态形状序列;在保留的静态配置下,平均形状误差为2.01厘米;动态轨迹跟踪在仿真与物理设备上均得到定性验证。每回合随机遮蔽30%的磁传感器,生成单一鲁棒策略,无需备用控制器或模式切换逻辑。采用非对称演员-评论家架构,利用特权平衡信息提升部分可观测下的价值估计;演员网络附加形状重建头,实现从原始诊断数据端到端重构形状,同时作为策略分析的可解释工具。该策略成功迁移至实际DIII-D放电实验中,直接控制线圈执行两次动态形状操作,并在独立的GSevolve仿真器中验证有效性。

原文摘要 · Abstract (English)

Plasma shape control in tokamaks requires a real-time controller that tracks dynamically changing shape targets while tolerating diagnostic failures. Classical approaches decompose the problem into equilibrium reconstruction followed by a linear controller, and assume a fixed, fully operational sensor set. We present a reinforcement learning agent that addresses both limitations simultaneously. The agent is trained in NSFsim, a high-fidelity tokamak simulator configured for DIII-D, on a curated dataset of 120 experimental plasma shapes. The shape targets are resampled as random step changes every 0.25 s, exposing the agent to diverse transitions across the full shape envelope. At test time the agent zero-shot tracks dynamic shape sequences; on a held-out static configuration in simulation it achieves a mean shape error of 2.01 cm, and dynamic trajectory following is demonstrated qualitatively in simulation and on the physical device. Diagnostic dropout randomly masks 30% of magnetic sensors per episode, yielding a single policy robust to arbitrary sensor subsets without backup controllers or mode-switching logic. An asymmetric actor-critic architecture with privileged equilibrium information improves value estimation under partial observability; an auxiliary shape reconstruction head on the actor enables end-to-end shape reconstruction from raw diagnostics and serves as an interpretability tool for policy analysis. The policy transfers to experimental DIII-D shots, where it directly commands the coil actuators on two dynamic shape maneuvers, and to the independent GSevolve simulator.

等离子体控制强化学习聚变能源动态控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。