arXiv:2608.04222physics.flu-dyncs.LG2026-08

TIDE构建了256³的3D湍流基准数据集,推动科学机器学习发展

TIDE: A Physically Diverse 3D Turbulence Benchmark Dataset for Advancing Scientific Machine Learning

论文配图:TIDE: A Physically Diverse 3D Turbulence Benchmark Dataset for Advancing Scientific Machine Learning
图 1 · 摘自论文原文
  • 构建15种配置、多独立样本的3D不可压缩湍流数据集
  • 当前模型预测误差是谱求解器的两倍,且点误差低未必物理真实
  • 适合研究物理信息机器学习、湍流建模与泛化能力的学者

湍流是机器学习模拟物理动力学的核心测试平台,因其控制方程完全已知。然而现有研究多集中于二维,而三维湍流具有根本不同的物理特性且模拟成本更高。现有3D数据集通常仅提供单个实现,难以区分模型学习的是动力学规律还是单一流动的统计特征。本文提出TIDE(Turbulent Incompressible DNS Ensembles),一个256³的直接数值模拟数据集和基准,涵盖15种配置、8个可控轴、独立样本集、压力场及方程级验证。基准包含五项任务、标准化学习基线、受控泛化划分及物理保真度指标。在主要预测配置中,当前学习模型性能仅略优于持续性预测,误差约为已知方程谱求解器的两倍。此外,点误差较低可能伴随小尺度动力学严重失真,表明精度不等于物理保真。泛化结果进一步显示,多数模式转移源于训练覆盖不足;而强制衰减迁移暴露了缺失的条件变量:在无外力驱动时,强制训练的算子仍预测受驱演化。填补这些精度、保真与条件性差距,正是由TIDE可量化的核心开放问题。

原文摘要 · Abstract (English)

Turbulence is a central testbed for machine learning on physical dynamics because its governing laws are known exactly. However, most existing studies remain in 2D, while 3D turbulence has fundamentally different physics and is far more costly to simulate. Existing 3D resources also typically provide only one realization per configuration, making it difficult to distinguish learning the dynamics from fitting the statistics of a single flow. In this paper, we introduce TIDE (Turbulent Incompressible DNS Ensembles), a 256^3 DNS corpus and benchmark for 3D incompressible turbulence, with 15 configurations on eight controlled axes, independent ensembles, pressure fields, and equation-level verification. The benchmark includes five tasks, standardized learned baselines, controlled generalization splits, and physical-fidelity metrics alongside pointwise error. Across the main forecasting configurations, current learned models barely outperform persistence and still make about twice the error of a spectral solver given the true equations. Moreover, lower pointwise error can coincide with severely distorted small-scale dynamics, showing that accuracy alone does not ensure physical fidelity. Generalization results further show that most regime shifts reflect limited training coverage, whereas forced-to-decay transfer exposes a missing conditioning variable: operators trained under forcing continue to predict driven evolution when the external drive is removed. Closing these accuracy, fidelity, and conditioning gaps is the central open problem made measurable by TIDE.

湍流建模科学机器学习物理信息网络基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。