arXiv:2605.21211eess.SYcs.LG2026-05中稿 · publication at the…

YANN-RL让化学过程控制更高效可靠,训练快、数据少、性能接近最优控制。

Reinforcement Learning-based Control via Y-wise Affine Neural Networks: Comparative Case Studies for Chemical Processes

论文配图:Reinforcement Learning-based Control via Y-wise Affine Neural Networks: Comparative Case Studies for Chemical Processes
图 1 · 摘自论文原文
  • 用特定初始化的神经网络加速强化学习收敛,提升可解释性
  • 三类化工案例中训练时间与数据量大幅减少,逼近非线性模型预测控制性能
  • 适合工业界部署,无需完整系统模型,可信度高

本文提出一种高效且可实际应用的强化学习(RL)控制方法,用于化学过程系统。由于难以信任RL算法以及训练可靠智能体耗时长,该领域尚未广泛采用基于RL的控制。为此,我们采用前期工作提出的Y-wise Affine Neural Network(YANN)-RL算法。通过战略性地初始化智能体的策略与价值网络,YANN-RL在控制框架中提供自信且可解释的起点。我们将该方法应用于三个公开于PC-Gym库的流程工程案例:连续搅拌釜反应器(CSTR)、四水箱系统和多级萃取塔。与PPO、SAC、DDPG和TD3等主流算法对比,并以非线性模型预测控制(NMPC)为基准。结果表明,YANN-RL显著降低训练时间和所需数据量,可在化学过程系统中可靠部署,并在无需完整非线性模型的情况下达到接近NMPC的性能。

原文摘要 · Abstract (English)

In this work we present an efficient and practically implementable approach for the application of reinforcement learning (RL)-based control in chemical process systems. This is an area that has yet to widely adopt RL-based control largely due to inherent challenges in trusting RL algorithms and the time-consuming process of training reliable agents. To address these challenges, we leverage a class of RL algorithms termed Y-wise Affine Neural Network (YANN)- RL, which we have developed in our prior work (Braniff and Tian, 2025a). By strategically initializing actor and critic networks YANN-RL algorithms provide confident and interpretable starting points within control schemes. We apply this RL-based control approach to three different process engineering case studies publicly available on the PC-Gym library (Bloor et al., 2026): (i) a continuous stirred tank reactor (CSTR), (ii) a four-tank system, and (iii) a multistage extraction column. Our approach is compared to several popular RL algorithms (PPO, SAC, DDPG, and TD3) and is benchmarked against nonlinear model predictive control (NMPC). These case studies demonstrate that YANN-RL can greatly reduce the training time and data needed, can be deployed with confidence for chemical process systems, and can approach the performance of NMPC without the knowledge of a full nonlinear model.

强化学习化工控制神经网络模型预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。