arXiv:2602.01379physics.flu-dyncs.LG2026-02

构建大规模高雷诺数湍流数据集,助力机器学习预测复杂水下流动。

WAKESET: A Large-Scale, High-Reynolds Number Flow Dataset for Machine Learning of Turbulent Wake Dynamics

  • 基于1091个高保真模拟,扩展至4364组数据,覆盖高速与转向工况。
  • 最高雷诺数达1.09×10⁸,捕捉真实工程场景下的多尺度湍流特征。
  • 适合从事流体预测、代理建模与水下自主导航的科研与工程人员。

机器学习有望革新计算流体动力学,实现加速仿真、改进湍流建模及实时流动预测与控制。然而,其发展受限于缺乏大规模、多样化且高保真的数据集,尤其在高雷诺数湍流领域更为突出。本文提出WAKESET,一个针对高湍流流动的大规模CFD数据集,聚焦于大型无人潜航器回收小型自主潜航器的水下过程。包含1,091个高保真雷诺平均纳维-斯托克斯(RANS)模拟,扩展至4,364个实例,覆盖速度范围高达雷诺数1.09×10⁸,以及多种转向角度。该数据集结构清晰,建模与验证流程完备,面向实际工程问题,具有显著规模与高湍流特性,可有效支持机器学习模型在流场预测、代理建模和水下自主导航等任务中的开发与基准测试。

原文摘要 · Abstract (English)

Machine learning (ML) offers transformative potential for computational fluid dynamics (CFD), promising to accelerate simulations, improve turbulence modelling, and enable real-time flow prediction and control-capabilities that could fundamentally change how engineers approach fluid dynamics problems. However, the exploration of ML in fluid dynamics is critically hampered by the scarcity of large, diverse, and high-fidelity datasets suitable for training robust models. This limitation is particularly acute for highly turbulent flows, which dominate practical engineering applications yet remain computationally prohibitive to simulate at scale. High-Reynolds number turbulent datasets are essential for ML models to learn the complex, multi-scale physics characteristic of real-world flows, enabling generalisation beyond the simplified, low-Reynolds number regimes often represented in existing datasets. This paper introduces WAKESET, a novel, large-scale CFD dataset of highly turbulent flows, designed to address this critical gap. The dataset captures the complex hydrodynamic interactions during the underwater recovery of an autonomous underwater vehicle by a larger extra-large uncrewed underwater vehicle. It comprises 1,091 high-fidelity Reynolds-Averaged Navier-Stokes simulations, augmented to 4,364 instances, covering a wide operational envelope of speeds (up to Reynolds numbers of 1.09 x 10^8) and turning angles. This work details the motivation for this new dataset by reviewing existing resources, outlines the hydrodynamic modelling and validation underpinning its creation, and describes its structure. The dataset's focus on a practical engineering problem, its scale, and its high turbulence characteristics make it a valuable resource for developing and benchmarking ML models for flow field prediction, surrogate modelling, and autonomous navigation in complex underwater environments.

流体模拟湍流建模数据集机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。