用机器学习优化污水脱氮,揭示数据质量与模型泛化关键因素
Machine learning in wastewater treatment: insights from modelling a pilot denitrification reactor
- 基于挪威试点反应器数据,探索脱氮过程的机器学习建模方法
- 非线性模型训练表现好,但线性模型在时间后段测试中更稳定
- 需多年数据应对气候波动,适合污水处理领域研究者参考
污水处理厂因社会重要性和丰富数据成为机器学习应用的潜力场景,但其设计多样、运行条件复杂及进水特性多变,制约了自动化。本研究利用挪威Veas处理厂试点反应器的数据,探讨机器学习在生物反硝化过程中将硝酸盐(NO₃⁻)还原为氮气(N₂)的优化应用。研究不只关注预测精度,更强调构建有效数据驱动模型的基础条件:关键过程参数识别、所需数据量与质量、数据结构设计及模型特性要求。结果表明,非线性模型在训练和验证集上表现更优,说明存在非线性关系;而线性模型在时间靠后的未见测试数据上泛化能力更强。水温变量因训练与测试数据分布显著差异,对模型性能有严重负面影响。因此,我们得出结论:需多年数据才能建立鲁棒的机器学习模型。该研究为北方地区受气候波动影响的污水处理提供结构化、定制化的机器学习应用基础。文中公开了全部数据与代码。
原文摘要 · Abstract (English)
Wastewater treatment plants are increasingly recognized as promising candidates for machine learning applications, due to their societal importance and high availability of data. However, their varied designs, operational conditions, and influent characteristics hinder straightforward automation. In this study, we use data from a pilot reactor at the Veas treatment facility in Norway to explore how machine learning can be used to optimize biological nitrate ($\mathrm{NO_3^-}$) reduction to molecular nitrogen ($\mathrm{N_2}$) in the biogeochemical process known as \textit{denitrification}. Rather than focusing solely on predictive accuracy, our approach prioritizes understanding the foundational requirements for effective data-driven modelling of wastewater treatment. Specifically, we aim to identify which process parameters are most critical, the necessary data quantity and quality, how to structure data effectively, and what properties are required by the models. We find that nonlinear models perform best on the training and validation data sets, indicating nonlinear relationships to be learned, but linear models transfer better to the unseen test data, which comes later in time. The variable measuring the water temperature has a particularly detrimental effect on the models, owing to a significant change in distributions between training and test data. We therefore conclude that multiple years of data is necessary to learn robust machine learning models. By addressing foundational elements, particularly in the context of the climatic variability faced by northern regions, this work lays the groundwork for a more structured and tailored approach to machine learning for wastewater treatment. We share publicly both the data and code used to produce the results in the paper.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。