arXiv:2601.15503cs.LG2026-01

用机器学习填补湖水监测数据空缺,高效预测水质变化。

Data-driven Lake Water Quality Forecasting for Time Series with Missing Data using Machine Learning

  • 用多重插补法处理缺失数据,选岭回归做预测。
  • 仅需约176个样本和4个特征,就能达95%准确率。
  • 提出联合可行性函数,指导监测采样与测项选择。

志愿性湖泊监测产生不规则、季节性的时序数据,常因冰盖、天气限制和人为失误导致大量缺失,给有害藻华的预测与预警带来挑战。本文基于缅因州30个湖泊三十年的实地观测数据,研究了透明度盘深度(SDD)的预测。采用多重插补法(MICE)处理缺失值,并以归一化平均绝对误差(nMAE)评估跨湖可比性能。在六种模型中,岭回归表现最佳。进一步分析表明,在采用后向近期历史策略下,平均每湖约176个训练样本即可达到完整历史数据95%的精度。同时识别出最小特征集:4个特征组合可实现与13个特征基准相差不超过5%的性能。据此提出联合可行性函数,统一确定满足5%误差目标所需的最少历史长度与最小特征数。结果显示,仅需约64个近期样本和每湖一个关键指标,即可达成目标,凸显针对性监测的可行性。该策略将近期历史长度与特征选择整合于固定精度目标下,为湖沼研究人员提供简洁高效的采样与测量优先级制定依据。

原文摘要 · Abstract (English)

Volunteer-led lake monitoring yields irregular, seasonal time series with many gaps arising from ice cover, weather-related access constraints, and occasional human errors, complicating forecasting and early warning of harmful algal blooms. We study Secchi Disk Depth (SDD) forecasting on a 30-lake, data-rich subset drawn from three decades of in-situ records collected across Maine lakes. Missingness is handled via Multiple Imputation by Chained Equations (MICE), and we evaluate performance with a normalized Mean Absolute Error (nMAE) metric for cross-lake comparability. Among six candidates, ridge regression provides the best mean test performance. Using ridge regression, we then quantify the minimal sample size, showing that under a backward, recent-history protocol, the model reaches within 5% of full-history accuracy with approximately 176 training samples per lake on average. We also identify a minimal feature set, where a compact four-feature subset matches the thirteen-feature baseline within the same 5% tolerance. Bringing these results together, we introduce a joint feasibility function that identifies the minimal training history and fewest predictors sufficient to achieve the target of staying within 5% of the complete-history, full-feature baseline. In our study, meeting the 5% accuracy target required about 64 recent samples and just one predictor per lake, highlighting the practicality of targeted monitoring. Hence, our joint feasibility strategy unifies recent-history length and feature choice under a fixed accuracy target, yielding a simple, efficient rule for setting sampling effort and measurement priorities for lake researchers.

水质预测缺失数据机器学习湖沼监测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。