对比三种AI模型在大规模水质预测中的可信度,发现关键变量预测最不准确且易受数据异常影响。
Identifying Trustworthiness Challenges in Deep Learning Models for Continental-Scale Water Quality Prediction
- 用三种深度学习模型在482个流域37年数据上评估水质预测可信度
- 复杂过程与数据稀缺导致关键水质变量预测误差大、不确定性高
- 模型对异常值敏感,跨流域泛化能力差,需多方法解释以提升可信度
水质关乎环境可持续性、生态韧性与公共健康。深度学习虽具大规模水质预测潜力,但因性能差异、鲁棒性、不确定性、可解释性、泛化性和可复现性等可信度问题,难以用于污染治理和资源公平分配等高风险决策。本文基于37年美国482个流域数据,对三种前沿深度学习模型(LSTM、DeepONet、Informer)进行多维度量化评估,预测20项水质指标。结果表明:水质变量的预测性能随过程复杂度与数据可得性变化,管理关键变量预测最差、不确定性最高;鲁棒性测试显示模型对异常值和目标污染极为敏感,其中基线表现最佳的LSTM在数据损坏时最为脆弱;可解释性分析中,简单变量的归因一致,但营养物变量分歧显著,凸显多方法解释必要性;所有模型在未观测流域的泛化能力均较差。本研究呼吁推进可信的智能水管理方法,并为研究人员、决策者和实践者提供负责任使用AI的路径。
原文摘要 · Abstract (English)
Water quality is foundational to environmental sustainability, ecosystem resilience, and public health. Deep learning offers transformative potential for large-scale water quality prediction and scientific insights generation. However, their widespread adoption in high-stakes operational decision-making, such as pollution mitigation and equitable resource allocation, is prevented by unresolved trustworthiness challenges, including performance disparity, robustness, uncertainty, interpretability, generalizability, and reproducibility. In this work, we present a multi-dimensional, quantitative evaluation of trustworthiness benchmarking three state-of-the-art deep learning architectures: recurrent (LSTM), operator-learning (DeepONet), and transformer-based (Informer), trained on 37 years of data from 482 U.S. basins to predict 20 water quality variables. Our investigation reveals systematic performance disparities tied to process complexity, data availability, and basin heterogeneity. Management-critical variables remain the least predictable and most uncertain. Robustness tests reveal pronounced sensitivity to outliers and corrupted targets; notably, the architecture with the strongest baseline performance (LSTM) proves most vulnerable under data corruption. Attribution analyses align for simple variables but diverge for nutrients, underscoring the need for multi-method interpretability. Spatial generalization to ungauged basins remains poor across all models. This work serves as a timely call to action for advancing trustworthy data-driven methods for water resources management and provides a pathway to offering critical insights for researchers, decision-makers, and practitioners seeking to leverage artificial intelligence (AI) responsibly in environmental management.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。