arXiv:2510.03830cs.LGcs.SY2025-10被引 1

用混合离线学习与在线优化提升工厂启停和换产控制,超越人类专家表现。

HOFLON: Hybrid Offline Learning and Online Optimization for Process Start-Up and Grade-Transition Control

  • 离线构建数据流形与长期价值函数,识别可行操作区域。
  • 在线优化时兼顾奖励最大化与变量变化平滑,避免越界风险。
  • 在两种工业场景中均优于历史最优,适合自动化生产系统应用。

连续流程工厂的启停和产品换产是关键操作环节,任何失误都会立即影响产品质量并造成运营损失。这些操作长期依赖少数资深操作员手动执行,但随着人员退休,相关隐性知识正逐渐流失。在缺乏过程模型的情况下,离线强化学习可通过挖掘历史启停与换产记录来捕捉甚至超越人类经验,但标准离线强化学习在策略超出数据范围时易受分布偏移和价值过估计影响。本文提出HOFLON(混合离线学习+在线优化),离线阶段学习两个核心组件:(i) 表征过去过渡行为可行区域的潜在数据流形,(ii) 能预测状态-动作对累计奖励的长时程Q评价器。在线阶段,求解一步优化问题,以最大化Q值为目标,同时惩罚偏离已学流形及操纵变量变化率过高的行为。我们在两个工业案例中验证:聚合反应器启停与纸机换产,与领先的离线强化学习算法IQL对比,HOFLON不仅全面超越其表现,且平均累积奖励高于历史最佳记录,证明其具备超越现有专家能力实现过渡操作自动化的潜力。

原文摘要 · Abstract (English)

Start-ups and product grade-changes are critical steps in continuous-process plant operation, because any misstep immediately affects product quality and drives operational losses. These transitions have long relied on manual operation by a handful of expert operators, but the progressive retirement of that workforce is leaving plant owners without the tacit know-how needed to execute them consistently. In the absence of a process model, offline reinforcement learning (RL) promises to capture and even surpass human expertise by mining historical start-up and grade-change logs, yet standard offline RL struggles with distribution shift and value-overestimation whenever a learned policy ventures outside the data envelope. We introduce HOFLON (Hybrid Offline Learning + Online Optimization) to overcome those limitations. Offline, HOFLON learns (i) a latent data manifold that represents the feasible region spanned by past transitions and (ii) a long-horizon Q-critic that predicts the cumulative reward from state-action pairs. Online, it solves a one-step optimization problem that maximizes the Q-critic while penalizing deviations from the learned manifold and excessive rates of change in the manipulated variables. We test HOFLON on two industrial case studies: a polymerization reactor start-up and a paper-machine grade-change problem, and benchmark it against Implicit Q-Learning (IQL), a leading offline-RL algorithm. In both plants HOFLON not only surpasses IQL but also delivers, on average, better cumulative rewards than the best start-up or grade-change observed in the historical data, demonstrating its potential to automate transition operations beyond current expert capability.

强化学习工业控制流程优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。