用强化学习优化化工操作配方,更安全、更可解释。
Optimizing Operation Recipes with Reinforcement Learning for Safe and Interpretable Control of Chemical Processes
- 基于专家经验的配方结构,用强化学习调参
- 仿真显示性能接近最优控制器,数据需求少
- 适合需要安全与可解释性的工业控制场景
化工过程的最优运行对节能、降耗和降本至关重要。传统强化学习面临质量与安全硬约束难以满足、训练数据量大的难题。化工过程实验数据有限,而详细动态模型虽可替代却因复杂性导致计算不可行;模型预测控制等方法也受制于模型复杂度。因此,多数工艺仍依赖人工制定的配方与简单线性控制器,导致性能不佳且灵活性差。本文提出新方法:利用嵌入在操作配方中的专家知识,通过强化学习优化配方参数及其底层线性控制器,生成优化后的操作策略。该方法显著减少数据需求,更有效处理约束,且因配方结构化而更具可解释性。在工业级间歇聚合反应器的仿真中验证,其性能可逼近最优控制器,同时克服了现有方法的局限。
原文摘要 · Abstract (English)
Optimal operation of chemical processes is vital for energy, resource, and cost savings in chemical engineering. The problem of optimal operation can be tackled with reinforcement learning, but traditional reinforcement learning methods face challenges due to hard constraints related to quality and safety that must be strictly satisfied, and the large amount of required training data. Chemical processes often cannot provide sufficient experimental data, and while detailed dynamic models can be an alternative, their complexity makes it computationally intractable to generate the needed data. Optimal control methods, such as model predictive control, also struggle with the complexity of the underlying dynamic models. Consequently, many chemical processes rely on manually defined operation recipes combined with simple linear controllers, leading to suboptimal performance and limited flexibility. In this work, we propose a novel approach that leverages expert knowledge embedded in operation recipes. By using reinforcement learning to optimize the parameters of these recipes and their underlying linear controllers, we achieve an optimized operation recipe. This method requires significantly less data, handles constraints more effectively, and is more interpretable than traditional reinforcement learning methods due to the structured nature of the recipes. We demonstrate the potential of our approach through simulation results of an industrial batch polymerization reactor, showing that it can approach the performance of optimal controllers while addressing the limitations of existing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。