用深度学习提升环境预测的精度、效率与可解释性,解决洪水、天气和科学问答难题。
Accurate, Efficient, and Explainable Deep Learning Approaches for Environmental Science Problems

- 设计轻量级水位预测模型与可解释的决策模型,实现沿海洪水实时管理
- 提出条件扩散模型CoDiCast,实现全球天气的概率化精准预测并量化不确定性
- 构建结构化知识框架的Hypercube-RAG,兼顾问答准确、高效与透明可追溯
环境科学在保护生态系统中至关重要,依赖大规模异构数据。在大数据时代,人工智能成为发现模式与支持决策的关键工具。本论文针对环境科学中的三类复杂问题,发展适配的AI方法以实现环境智能。首先,在佛罗里达南部沿海洪涝易发区,针对传统物理模型计算成本高、难以实时应用的问题,提出深度学习模型WaLeF进行水位预测,并设计可解释的预报驱动模型FIDLAr,显著提升预测精度与效率。其次,针对全球天气预测面临的海量数据挑战,提出条件扩散模型CoDiCast,基于生成式AI实现概率化预报,兼具准确性、高效性与明确的不确定性量化能力。最后,针对大语言模型在环境科学问答中常出现幻觉的问题,提出基于结构化文本立方体框架的Hypercube-RAG,克服现有检索增强生成方法在准确性、效率与可解释性之间的权衡,实现三者兼备。
原文摘要 · Abstract (English)
Environmental science plays a pivotal role in safeguarding ecosystems, a domain driven by large-scale, heterogeneous data. In the big data era, artificial intelligence (AI) has emerged as a transformative tool for learning patterns and supporting decision-making. This dissertation develops AI-based approaches tailored to complex environmental science problems to achieve Environmental Intelligence, studying three specific challenges. First, we focus on flood prediction and management in coastal river systems. Conventional physics-based models are computationally intensive, limiting real-time application. To overcome this, we propose a deep learning (DL)-based model, WaLeF, for water level forecasting, and a forecast-informed DL model, FIDLAr, to manage water levels. Evaluated in a flood-prone coastal system in South Florida characterized by extreme rainfall and sea level fluctuations, FIDLAr outperforms baselines in accuracy and efficiency while providing interpretable outputs. Second, we target global weather prediction, which is challenged by massive data scale. Traditional physics methods are deterministic and computationally heavy. We propose CoDiCast, a conditional diffusion model tailored for probabilistic weather forecasting. Adapted from generative AI for predictive tasks, experiments show CoDiCast achieves accurate, efficient forecasts with explicit uncertainty quantification. Lastly, we address scientific question-answering in environmental science. When answering in-domain questions, large language models (LLMs) often suffer from hallucinations due to out-of-date or limited knowledge. While retrieval-augmented generation (RAG) retrieves domain-specific knowledge, existing methods trade off accuracy, efficiency, or explainability. We propose Hypercube-RAG, built on a structured text cube framework, which successfully exhibits all three properties simultaneously.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。