arXiv:2509.06853eess.SYcs.AI2025-09被引 12

用强化学习+行为克隆控制光生物反应器pH值,实测稳定高效。

Reinforcement learning meets bioprocess control through behaviour cloning: Real-world deployment in an industrial photobioreactor

  • 先离线学PID轨迹,再在线微调,适应环境变化。
  • 误差降8%,控制能耗比PID低54%,抗干扰更强。
  • 首次在真实开放反应器中应用,适合工业生物过程控制。

活细胞作为生产单元具有固有复杂性,尤其在受环境波动影响的开放式光生物反应器(PBR)中维持稳定最优工艺条件面临巨大挑战。为此,本文提出一种结合强化学习(RL)与行为克隆(BC)的pH调控方法,据我们所知,这是首个将基于RL的控制策略应用于此类非线性且易受扰动的生物过程。方法包括离线训练阶段:RL代理从由标准比例-积分-微分(PID)控制器生成的轨迹中学习,不直接交互真实系统;随后每日在线微调,以适应过程动态变化并有效抑制快速瞬态扰动。该混合离线-在线策略使自适应控制策略能应对开放PBR中的固有非线性及外部扰动。仿真结果表明,相比PID控制,积分绝对误差(IAE)降低8%;相比标准离线策略,降低5%。同时控制努力显著减少——较PID降低54%,较标准RL降低7%,对降低运行成本至关重要。最后,在多变环境条件下进行了为期8天的实验验证,确认了方法的鲁棒性和可靠性。本工作展示了基于强化学习的生物过程控制潜力,并为其他非线性、易受扰系统提供可推广路径。

原文摘要 · Abstract (English)

The inherent complexity of living cells as production units creates major challenges for maintaining stable and optimal bioprocess conditions, especially in open Photobioreactors (PBRs) exposed to fluctuating environments. To address this, we propose a Reinforcement Learning (RL) control approach, combined with Behavior Cloning (BC), for pH regulation in open PBR systems. This represents, to the best of our knowledge, the first application of an RL-based control strategy to such a nonlinear and disturbance-prone bioprocess. Our method begins with an offline training stage in which the RL agent learns from trajectories generated by a nominal Proportional-Integral-Derivative (PID) controller, without direct interaction with the real system. This is followed by a daily online fine-tuning phase, enabling adaptation to evolving process dynamics and stronger rejection of fast, transient disturbances. This hybrid offline-online strategy allows deployment of an adaptive control policy capable of handling the inherent nonlinearities and external perturbations in open PBRs. Simulation studies highlight the advantages of our method: the Integral of Absolute Error (IAE) was reduced by 8% compared to PID control and by 5% relative to standard off-policy RL. Moreover, control effort decreased substantially-by 54% compared to PID and 7% compared to standard RL-an important factor for minimizing operational costs. Finally, an 8-day experimental validation under varying environmental conditions confirmed the robustness and reliability of the proposed approach. Overall, this work demonstrates the potential of RL-based methods for bioprocess control and paves the way for their broader application to other nonlinear, disturbance-prone systems.

强化学习生物过程控制工业自动化行为克隆

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。