arXiv:2608.23885eess.SYcs.AI2026-08

数据驱动模型虽拟合精准,却可能产生虚假最优解,误导实时优化。

A tale of perfect fit and phantom optima: how data-driven models can fail in real-time optimization

论文配图:A tale of perfect fit and phantom optima: how data-driven models can fail in real-time optimization
图 1 · 摘自论文原文
  • 融合机理与神经网络的混合建模,提升对复杂工艺的逼近能力。
  • 模型拟合误差小,但优化结果与真实工厂最优解差异显著。
  • 适合关注工业优化可靠性的研究人员和工程师参考。

实时优化(RTO)依赖过程模型寻找经济最优操作条件。由于基于机理的模型需大量工艺知识,数据驱动方法日益受青睐。现代机器学习模型能精确拟合历史生产数据,并通过常规验证测试。然而,这类模型是否适用于经济优化仍不明确。本文以醋酸乙烯酯基准工艺为例,该工艺具有唯一且良好条件的经济最优解。我们训练了结合已知质量平衡与热力学约束的结构化混合模型,以及全数据驱动的神经微分方程(Neural ODE)模型。两者均能准确再现工厂测量数据,且在不同随机初始化下预测结果变化极小。然而,它们的经济最优解与实际工厂结果存在显著差异:工厂在多起点搜索中仅返回单一最优解,而训练模型却出现多个虚假最优解。此外,我们发现训练优化器本身也可能是误差来源——即使使用无噪声数据,初始权重接近真实最优解,随机梯度训练仍可能漂移至导致严重性能下降的权重。因此,所识别的模型不仅是数据的产物,也是训练优化器的副产品。这些结果表明,对所有可用测量值的良好预测拟合,并不能保证可靠的经济性能。用于RTO的数据驱动模型,至少应在决策导向的基准测试(如本文开发的基准)中恢复真实最优解后,才可考虑应用于实际工厂。

原文摘要 · Abstract (English)

Real-time optimization (RTO) relies on process models to locate economically optimal operating conditions. Because developing first-principles models requires significant process knowledge, data-driven alternatives are increasingly attractive. Modern machine-learning models can fit historical plant data accurately and often pass standard validation tests. Whether such models can be trusted for economic optimization, however, remains unclear. We investigate this question using a vinyl acetate monomer benchmark process with a unique, well-conditioned economic optimum. We train a structured hybrid model that combines known mass balances and thermodynamics with a neural-network closure for unknown kinetics, and a fully data-driven neural ordinary differential equation (ODE) model. Both models reproduce plant measurements accurately and exhibit little variation in predictions across random initializations. Yet their economic optima differ substantially from that of the plant. Where the plant returns a single optimum on multistart search, the trained models return many phantom optima. We further show that the training optimizer alone can be yet another source of error. Even with noise-free data and initialization at weights that recover the plant optimum, stochastic gradient training can drift to weights that yield substantially worse RTO solutions. The identified model is thus an artifact of the training optimizer as well as the data. These results demonstrate that a good predictive fit of all available measurements does not guarantee reliable economic performance. A data-driven model for RTO should at least be required to recover the optimum on a decision-oriented benchmark like the one developed here before being considered for plant testing and application.

实时优化数据驱动模型失效最优解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。