arXiv:2510.17531physics.plasm-phcs.LG2025-10被引 1

用历史数据训练零样本控制策略,实现等离子体形状的通用精准调控。

Plasma Shape Control via Zero-shot Generative Reinforcement Learning

  • 基于生成对抗模仿学习与希尔伯特空间表征,融合稳定控制风格与几何结构潜空间。
  • 在HL-3托卡马克仿真中,对多种等离子体场景实现零样本轨迹跟踪,精度高且稳定。
  • 适合未来聚变反应堆智能控制系统研发,无需任务微调即可部署。

传统PID控制器在等离子体形状控制中适应性有限,而任务特定的强化学习方法存在泛化能力差、需重复训练的问题。本文提出一种新框架,从大规模历史PID控制放电数据中构建通用的零样本控制策略。方法结合生成对抗模仿学习(GAIL)与希尔伯特空间表示学习,实现双重目标:模仿PID数据的稳定运行风格,并构建几何结构化的潜在空间以支持高效、目标导向的控制。所得到的基础策略可在无任务微调的情况下,直接应用于多种轨迹跟踪任务。在HL-3托卡马克模拟器上的评估表明,该策略在不同等离子体场景下均能精确、稳定地跟踪关键形状参数参考轨迹。本工作为未来聚变反应堆开发高度灵活且数据高效的智能控制系统提供了可行路径。

原文摘要 · Abstract (English)

Traditional PID controllers have limited adaptability for plasma shape control, and task-specific reinforcement learning (RL) methods suffer from limited generalization and the need for repetitive retraining. To overcome these challenges, this paper proposes a novel framework for developing a versatile, zero-shot control policy from a large-scale offline dataset of historical PID-controlled discharges. Our approach synergistically combines Generative Adversarial Imitation Learning (GAIL) with Hilbert space representation learning to achieve dual objectives: mimicking the stable operational style of the PID data and constructing a geometrically structured latent space for efficient, goal-directed control. The resulting foundation policy can be deployed for diverse trajectory tracking tasks in a zero-shot manner without any task-specific fine-tuning. Evaluations on the HL-3 tokamak simulator demonstrate that the policy excels at precisely and stably tracking reference trajectories for key shape parameters across a range of plasma scenarios. This work presents a viable pathway toward developing highly flexible and data-efficient intelligent control systems for future fusion reactors.

等离子体控制生成模型强化学习零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。