arXiv:2605.00056cs.LGcs.AI2026-05中稿 · publication in Ear…

用变换+集成学习精准预测地下水重金属污染,结果更可靠。

Smart Ensemble Learning Framework for Predicting Groundwater Heavy Metal Pollution

论文配图:Smart Ensemble Learning Framework for Predicting Groundwater Heavy Metal Pollution
图 1 · 摘自论文原文
  • 对污染指数做对数或高斯耦合变换,再用集成模型预测。
  • 耦合模型达R²=0.96,RMSE=0.19,比原始模型更稳定可信。
  • 适合关注地下水污染评估与可解释性建模的研究者。

多瑙河流域地下水重金属污染日益严重,但传统方法难以捕捉污染指标的统计复杂性和空间异质性。关键挑战在于建模具有偏态分布且受相关污染物影响的重金属污染指数(HPI),未经变换会导致预测偏差。本研究构建了结合响应变换与嵌套交叉验证的集成学习框架。对HPI采用原始值、对数和高斯耦合三种变换,在六种学习器(支持向量回归、k-近邻、CART、弹性网络、核岭回归、堆叠Lasso集成)上进行评估。原始尺度模型呈现虚高拟合(弹性网络与堆叠集成的R²≈1.0),显示过度乐观。对数变换稳定方差(SVM: R²=0.93, RMSE=0.18;k-NN: R²=0.92, RMSE=0.20)。高斯耦合效果最佳:堆叠集成模型R²=0.96(RMSE=0.19),其余模型也保持高精度。耦合模型改善残差分布,生成空间合理地图。DBSCAN聚类识别出铁(Fe)和锰(Mn)为主要贡献因子,符合区域水文地球化学特征。局限包括依赖随机而非空间交叉验证,且研究范围限于该流域。未来工作应探索空间验证及其他地质环境。总体而言,基于分布感知的集成模型结合聚类诊断,提供鲁棒且可解释的地下水污染评估方案。

原文摘要 · Abstract (English)

Groundwater in the Densu Basin is increasingly threatened by heavy metal contamination, but conventional methods fail to capture the statistical complexity and spatial heterogeneity of pollution indicators. A key challenge is modelling the Heavy Metal Pollution Index (HPI), which is typically skewed and affected by correlated contaminants, leading to biased predictions without transformation. This study develops a predictive framework integrating response transformations with nested cross-validated ensemble machine learning. Three transformations (raw, log, and Gaussian copula) were applied to HPI and evaluated across six learners: support vector regression (SVM), $k$-nearest neighbours (k-NN), CART, Elastic Net, kernel ridge regression, and a stacked Lasso ensemble. Raw-scale models produced deceptively high fits (Elastic Net and stacked ensemble $R^2 \approx 1.0$), suggesting over-optimism. The log transformation stabilised variance (SVM: $R^2 = 0.93$, RMSE $= 0.18$; k-NN: $R^2 = 0.92$, RMSE $= 0.20$). The Gaussian copula gave the most reliable results: stacked ensemble $R^2 = 0.96$ (RMSE $= 0.19$), with other learners maintaining high accuracy. Copula-based models improved residuals and produced spatially plausible maps. DBSCAN clustering revealed Fe and Mn as primary HPI contributors, consistent with regional hydrogeochemistry. Limitations include reliance on random (not spatial) cross-validation and basin-specific scope. Future work should explore spatial validation and other geological settings. Overall, distribution-aware ensembles with clustering diagnostics offer robust, interpretable assessments of groundwater contamination.

污染预测集成学习地下水高斯耦合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。