arXiv:2504.00156eess.SYcs.LG2025-04被引 2

用深度强化学习控制核微堆,比传统方法更稳、更省力。

Nuclear Microreactor Control with Deep Reinforcement Learning

  • 用深度强化学习实现微堆实时棒控,支持多棒独立调节
  • 短时负载跟踪误差低于传统PID,长时运行误差在1%以内
  • 对噪声鲁棒性强,控制能耗更低,适合复杂场景

核微堆的经济可行性依赖于通过自主控制降低运行成本,尤其在与可再生能源协同运行时。本研究探索了深度强化学习(RL)在微堆实时棒控中的应用,重点评估其在负载跟踪场景下的表现。基于带热反馈和氙反馈的点堆模型,首先建立单输出RL代理的基准,再与传统比例-积分-微分(PID)控制器对比。结果显示,无论是单智能体还是多智能体强化学习(MARL),RL控制器在多种负载跟踪场景下均达到或超过PID性能。短期瞬态中,RL显著降低跟踪误差;在持续300分钟的长时负载跟踪中,虽氙反馈主导时PID精度略优,但RL仍保持在1%误差范围内,证明其强泛化能力,且仅需短时训练数据,避免过拟合。当控制扩展至多棒时,MARL实现各棒独立控制并自动满足反应堆对称性约束,而单智能体RL无法达成此目标。此外,在加入高斯噪声后,RL控制器仍保持更低误差与更小控制努力,优于PID。

原文摘要 · Abstract (English)

The economic feasibility of nuclear microreactors will depend on minimizing operating costs through advancements in autonomous control, especially when these microreactors are operating alongside other types of energy systems (e.g., renewable energy). This study explores the application of deep reinforcement learning (RL) for real-time drum control in microreactors, exploring performance in regard to load-following scenarios. By leveraging a point kinetics model with thermal and xenon feedback, we first establish a baseline using a single-output RL agent, then compare it against a traditional proportional-integral-derivative (PID) controller. This study demonstrates that RL controllers, including both single- and multi-agent RL (MARL) frameworks, can achieve similar or even superior load-following performance as traditional PID control across a range of load-following scenarios. In short transients, the RL agent was able to reduce the tracking error rate in comparison to PID. Over extended 300-minute load-following scenarios in which xenon feedback becomes a dominant factor, PID maintained better accuracy, but RL still remained within a 1% error margin despite being trained only on short-duration scenarios. This highlights RL's strong ability to generalize and extrapolate to longer, more complex transients, affording substantial reductions in training costs and reduced overfitting. Furthermore, when control was extended to multiple drums, MARL enabled independent drum control as well as maintained reactor symmetry constraints without sacrificing performance -- an objective that standard single-agent RL could not learn. We also found that, as increasing levels of Gaussian noise were added to the power measurements, the RL controllers were able to maintain lower error rates than PID, and to do so with less control effort.

强化学习核反应堆控制优化多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。