arXiv:2607.23983physics.geo-phcs.LG2026-07

将洪水预报员的隐性经验转化为可审计的工作流,提升预报准确性与可复现性。

HydroAgent: Formalizing Forecaster Expertise into Skill-Orchestrated Flood Forecasting Workflows

论文配图:HydroAgent: Formalizing Forecaster Expertise into Skill-Orchestrated Flood Forecasting Workflows
图 1 · 摘自论文原文
  • 用显式规则约束大模型推理,构建技能协调的洪水预报工作流
  • 14次洪峰事件中10次、11次分别在5%误差内捕捉峰值流量和洪水总量
  • 适合需提升预报透明度与决策支持的水利部门或灾害预警系统

业务化洪水预报依赖难以形式化、审计和转移的隐性预报员经验。尽管人工智能在洪水预测与模型误差修正方面取得进展,但现有研究大多未显式表达连接模型输出与预警决策的专家规则、审查节点与流程约束。为此,我们提出HydroAgent——一种将大语言模型(LLMs)嵌入模型驱动洪水预报流程的技能协调代理框架,每个技能编码明确规则以限定LLM推理范围。在南亚姆希尔河流域验证了五个先进LLMs的有效性。结果表明,在14次事件中,先验判断在5%容差内捕捉到10次峰值流量和11次洪水总量;129个事件经五折交叉验证,皮尔逊相关系数达0.62和0.84。基于高基线方案库(平均KGE 0.890),引导方案选择使KGE进一步提升0.023–0.154,模拟峰值流量与洪水总量在14和13次事件中均落在先验判断范围内。所有五种测试的LLMs均成功执行工作流,判断准确率相当(40%–80%),性能差异适中,成本差异显著。HydroAgent不旨在替代人类预报员,而是将隐性经验转化为可审计、可复现的工作流,精简分析步骤,支持更明智决策。该技能协调范式展示了显式规则边界如何引导语言模型推理,以补充物理模拟,推动下一代洪水预报发展。

原文摘要 · Abstract (English)

Operational flood forecasting depends on tacit forecaster expertise that is difficult to formalize, audit, and transfer. Although artificial intelligence methods have advanced flood prediction and model-error correction, most existing studies have not explicitly represented the tacit expert rules, review checkpoints, and workflow constraints that connect model outputs to operational warning decisions. To address this issue, we propose HydroAgent, a skill-orchestrated agent framework that embeds Large Language Models (LLMs) into a model-driven flood forecasting workflow, where each skill encodes explicit rules to bound LLM reasoning. We validated its effectiveness using five state-of-the-art LLMs in the South Yamhill River basin. Our results demonstrate that prior judgment captures observed peak flow and flood volume within 5% tolerance in 10 and 11 out of 14 events, with 5-fold cross-validation over 129 events yielding Pearson correlations of 0.62 and 0.84. Building on a high-baseline scheme library (average KGE 0.890), the guided scheme selection further improves KGE by 0.023-0.154, with simulated peak flow and flood volume falling within the prior judgment ranges for 14 and 13 out of 14 events. All five tested LLMs successfully execute the HydroAgent workflow with comparable judgment accuracy (40%-80%), while showing moderate performance variation and substantial cost differences. HydroAgent does not aim to replace human forecasters; instead, it translates their tacit expertise into an auditable and reproducible workflow, streamlining analytical steps and supporting more informed decision-making. This skill-orchestrated paradigm demonstrates how explicit rule boundaries can guide language model reasoning to complement physically based simulation in next-generation flood forecasting.

洪水预报大模型应用工作流自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。