arXiv:2604.09166cs.LG2026-04

构建大规模混合数据集,助力化工异常检测模型训练。

Automated Batch Distillation Process Simulation for a Large Hybrid Dataset for Deep Anomaly Detection

  • 用自动化流程生成仿真数据,结合实验数据构建混合数据集。
  • 仿真结果与真实实验高度一致,能准确复现正常及多种异常工况。
  • 适合从事化工过程监控、深度异常检测的研究者使用。

基于深度学习的化工过程异常检测潜力巨大,但缺乏大规模、多样化且标注完整的训练数据。此前我们构建了一个涵盖正常与异常工况的大规模全标注实验数据集。本文通过自动化工作流,利用新型基于Python的工艺模拟器,采用定制化的索引降维策略求解微分代数方程组,生成对应仿真数据,形成新颖的混合数据集。依托实验数据库丰富的元数据和结构化异常标注,实验记录被自动转换为仿真场景。经单个参考实验校准后,其余实验动态预测效果良好。实现了对大量实验运行的全自动化、一致性时间序列数据生成,覆盖正常操作及多种执行器与控制相关的异常。最终发布的混合数据集公开可用。从工艺模拟角度看,本工作展示了大规模实验周期的自动化、一致化仿真能力,以间歇精馏为例;从数据驱动异常检测角度,该混合数据集为仿真到实验的迁移学习、伪实验数据生成以及未来深度异常检测方法研究提供了独特基础。

原文摘要 · Abstract (English)

Anomaly detection (AD) in chemical processes based on deep learning offers significant opportunities but requires large, diverse, and well-annotated training datasets that are rarely available from industrial operations. In a recent work, we introduced a large, fully annotated experimental dataset for batch distillation under normal and anomalous operating conditions. In the present study, we augment this dataset with a corresponding simulation dataset, creating a novel hybrid dataset. The simulation data is generated in an automated workflow with a novel Python-based process simulator that employs a tailored index-reduction strategy for the underlying differential-algebraic equations. Leveraging the rich metadata and structured anomaly annotations of the experimental database, experimental records are automatically translated into simulation scenarios. After calibration to a single reference experiment, the dynamics of the other experiments are well predicted. This enabled the fully automated, consistent generation of time-series data for a large number of experimental runs, covering both normal operation and a wide range of actuator- and control-related anomalies. The resulting hybrid dataset is released openly. From a process simulation perspective, this work demonstrates the automated, consistent simulation of large-scale experimental campaigns, using batch distillation as an example. From a data-driven AD perspective, the hybrid dataset provides a unique basis for simulation-to-experiment style transfer, the generation of pseudo-experimental data, and future research on deep AD methods in chemical process monitoring.

异常检测工艺仿真混合数据集化工过程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。