arXiv:2603.15180eess.SYcs.AI2026-03

用迭代学习控制增强强化学习,让工业批处理更安全稳定

Iterative Learning Control-Informed Reinforcement Learning for Batch Process Control

  • 将卡尔曼滤波融入迭代学习框架,引导强化学习生成稳定策略
  • 在批间与批内双层控制中实现约束满足与收敛性保障
  • 适合需高可靠性、抗干扰的工业过程控制场景

深度强化学习(DRL)在探索-利用过程中存在动作的随机不确定性,带来训练和部署阶段的安全风险。工业过程控制中缺乏形式化的稳定性与收敛性保证,进一步限制了DRL的实际应用。相反,迭代学习控制(ILC)是针对重复性系统的成熟自主控制方法,尤其适用于批处理优化。ILC通过批次间或单批次内的控制律迭代修正,补偿重复性和非重复性扰动。本文提出一种迭代学习控制启发的强化学习框架(IL-CIRL),用于在批间与批内双层控制架构下训练DRL控制器。该方法在迭代学习结构中引入基于卡尔曼滤波的状态估计,引导DRL代理学习满足操作约束且具备稳定性保证的控制策略。该方法支持在多种扰动条件下系统化设计用于批处理过程的DRL控制器。

原文摘要 · Abstract (English)

A significant limitation of Deep Reinforcement Learning (DRL) is the stochastic uncertainty in actions generated during exploration-exploitation, which poses substantial safety risks during both training and deployment. In industrial process control, the lack of formal stability and convergence guarantees further inhibits adoption of DRL methods by practitioners. Conversely, Iterative Learning Control (ILC) represents a well-established autonomous control methodology for repetitive systems, particularly in batch process optimization. ILC achieves desired control performance through iterative refinement of control laws, either between consecutive batches or within individual batches, to compensate for both repetitive and non-repetitive disturbances. This study introduces an Iterative Learning Control-Informed Reinforcement Learning (IL-CIRL) framework for training DRL controllers in dual-layer batch-to-batch and within-batch control architectures for batch processes. The proposed method incorporates Kalman filter-based state estimation within the iterative learning structure to guide DRL agents toward control policies that satisfy operational constraints and ensure stability guarantees. This approach enables the systematic design of DRL controllers for batch processes operating under multiple disturbance conditions.

强化学习工业控制迭代学习稳定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。