arXiv:2605.17792cs.LGphysics.geo-ph2026-05被引 1

用模拟器反馈强化学习,让AI像专家一样精准校准水文模型。

HydroAgent: Closing the Gap Between Frontier LLMs and Human Experts in Hydrologic Model Calibration via Simulator-Grounded RL

论文配图:HydroAgent: Closing the Gap Between Frontier LLMs and Human Experts in Hydrologic Model Calibration via Simulator-Grounded RL
图 1 · 摘自论文原文
  • 基于仿真反馈的强化学习,让AI通过试错优化水文参数。
  • 最佳模型在4个流域上达到NSE 0.75,接近人类专家水平。
  • 适合需要高精度物理建模的气候与水资源研究者。

分布式水文模型校准是水资源管理中的关键瓶颈——径流预测、水库调度、干旱监测、基础设施设计和洪水预报均依赖于此。每个流域都需要专家将水文曲线特征转化为高维参数调整,且流程无法跨流域迁移。我们评估了九种前沿大语言模型代理(Claude Opus 4.6/4.7,Sonnet 4.6,GPT-5/5.4/5.4-pro,Gemini 2.5-pro/3.1-pro/3-flash)在美国家气象局用于洪水预报的CREST分布式水文模型上的表现。在四个独立测站(流域面积329–40,792 km²)上,最佳模型的二十轮平均纳什-萨特克利夫效率(NSE)从-0.16(GPT-5.4)到0.75(Sonnet 4.6),各厂商和能力层级均呈现类似上限,最强模型集中于0.65–0.75区间,仅Opus-4.7在单个站点达到人类专家水平。我们认为该差距并非由参数量决定,而是领域知识缺失所致。为此提出HYDROAGENT:以2,576条专家校准轨迹对Qwen3-4B进行监督微调,并采用组相对策略优化(Group-Relative Policy Optimization),利用在线CREST模拟输出的NSE作为可验证奖励信号,实现模拟器内闭环强化学习(RLSF)。对于地球系统科学,小规模领域微调策略配合模拟器驱动的强化学习,比盲目扩大通用大模型更具计算效率与物理一致性,而地球数据的多模态特性(遥感、原位时间序列、预报员文本)更凸显领域专用智能体的发展潜力。

原文摘要 · Abstract (English)

Calibrating distributed hydrologic models is a critical bottleneck across operational water resources management - streamflow prediction, reservoir operation, drought monitoring, infrastructure design, and flood forecasting all depend on it. Each basin demands an expert to translate hydrograph signatures into adjustments of a high-dimensional parameter vector, and the resulting workflow does not transfer between watersheds. We ask: can frontier large language model (LLM) agents replace the human hydrologic modeler, and if not, what would it take? We benchmark nine frontier LLM agents - Claude Opus 4.6/4.7, Sonnet 4.6, GPT-5/5.4/5.4-pro, and Gemini 2.5-pro/3.1-pro/3-flash - on the operational CREST distributed hydrologic model used by the U.S. National Weather Service for flash-flood forecasting. Best-of-twenty-rounds Nash-Sutcliffe Efficiency (NSE) across four held-out gauges spanning 329-40,792 km2 ranges from -0.16 (GPT-5.4) to 0.75 (Sonnet 4.6); the ceiling reproduces across all three vendors and capability tiers, with the strongest models concentrating in the 0.65-0.75 band, and no model reaches the human-expert reference except Opus-4.7 on one gauge. We argue this gap is not a parameter-count problem but a domain-grounding problem. We then propose HYDROAGENT, fine-tuning open-weight Qwen3-4B with supervised fine-tuning on 2,576 expert calibration trajectories and Group-Relative Policy Optimization using NSE as a verifiable reward from online CREST simulations - reinforcement learning with simulation feedback (RLSF). For Earth system science, a small domain-tuned policy with simulator-in-the-loop RL is a more compute-efficient and physically faithful path than scaling generic frontier models, and the multi-modal richness of Earth data - remote sensing, in-situ time series, and forecaster narrative - makes domain agents a leveraged direction for AI in physical science.

水文建模强化学习仿真反馈领域智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。