构建3.6万组水网仿真场景,助力机器学习研究
DiTEC-WDN: A Large-Scale Dataset of Hydraulic Scenarios across Multiple Water Distribution Networks
- 用自动化流程生成36,000个水网运行场景
- 涵盖22800万条图结构状态数据,支持多粒度建模
- 公开可用,解决真实水网数据隐私难题
隐私限制阻碍了真实供水管网(WDN)模型的共享,制约了数据驱动机器学习的应用,后者通常需要大量观测数据。为应对这一挑战,我们提出数据集DiTEC-WDN,包含在短周期(24小时)和长周期(1年)下模拟的36,000个独特场景。该数据集通过自动化管道构建,优化关键参数(如压力、流量、需求模式),实现大规模仿真,并通过规则验证与事后分析,记录符合水力实际的离散合成状态。共生成2.28亿条基于图的系统状态,可用于图级、节点级、边级回归及时间序列预测等多样化的机器学习任务。本成果以公开许可发布,推动水领域开放科研,消除敏感数据暴露风险,满足大规模供水网络基准研究与情景分析的需求。
原文摘要 · Abstract (English)
Privacy restrictions hinder the sharing of real-world Water Distribution Network (WDN) models, limiting the application of emerging data-driven machine learning, which typically requires extensive observations. To address this challenge, we propose the dataset DiTEC-WDN that comprises 36,000 unique scenarios simulated over either short-term (24 hours) or long-term (1 year) periods. We constructed this dataset using an automated pipeline that optimizes crucial parameters (e.g., pressure, flow rate, and demand patterns), facilitates large-scale simulations, and records discrete, synthetic but hydraulically realistic states under standard conditions via rule validation and post-hoc analysis. With a total of 228 million generated graph-based states, DiTEC-WDN can support a variety of machine-learning tasks, including graph-level, node-level, and link-level regression, as well as time-series forecasting. This contribution, released under a public license, encourages open scientific research in the critical water sector, eliminates the risk of exposing sensitive data, and fulfills the need for a large-scale water distribution network benchmark for study comparisons and scenario analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。