arXiv:2606.05265cs.LG2026-06

用少量数据实现跨流域洪水深度预测,无需每流域单独训练。

Data-efficient flood depth prediction through domain-aware coreset selection and tabular foundation models

论文配图:Data-efficient flood depth prediction through domain-aware coreset selection and tabular foundation models
图 1 · 摘自论文原文
  • 按暴雨周期和影响流域分层,智能选点构建核心数据集。
  • 仅用0.7%训练数据,平均准确率达0.663,接近全量数据的98.5%。
  • 适合需要快速部署、数据稀缺的洪水预警系统使用。

近实时洪水深度预测需要高效、准确且可跨流域迁移的代理模型。监督式代理模型虽能达到物理模拟器的精度,但需每个流域数百万条训练数据,且无法外推至原始网格范围之外。本文提出一种领域感知的核心数据集构建流程,在推理时对表格基础模型进行条件化。该流程按重现期和受影响流域分层,再通过目标感知的空间选择器采样六边形区域。仅使用每流域0.7%的训练数据,模型在休斯顿地区九个流域上的平均R²达到0.663,为全量监督参考模型(R²=0.673)的98.5%。模型可在未见流域上直接迁移,无需任务特定重训练;在真实暴雨事件中,于一个严重分布外案例上超越参考模型,而在一个大部分分布内案例上略逊一筹。领域感知的核心数据构造使表格基础模型无需每流域训练即可实现数据高效、流域可迁移的洪水预测。

原文摘要 · Abstract (English)

Near-real-time flood depth prediction demands surrogate models that are accurate, fast, and transferable across watersheds. Supervised surrogates can match physics-based simulators in accuracy but need millions of training rows per watershed and cannot extrapolate beyond their original mesh. We propose a domain-aware coreset construction pipeline that conditions a tabular foundation model at inference time. The pipeline stratifies storms by return period and most-affected watershed, then samples hexagons with a target-aware spatial selector. With 0.7% of the per-watershed training pool, the model attains a mean $R^2$ of 0.663 across nine Houston-area watersheds, within 98.5% of the supervised reference ($R^2$ = 0.673). It transfers to held-out watersheds without task-specific retraining, staying ahead of a coreset-trained supervised baseline. On real storms it exceeds the supervised reference on a far out-of-distribution case and trails it on a mostly in-distribution one. Domain-aware coreset construction lets tabular foundation models deliver data-efficient, watershed-transferable flood predictions without per-watershed training.

洪水预测数据效率迁移学习核心数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。