arXiv:2412.00568cs.LGphysics.flu-dyn2024-12NeurIPS被引 162

构建15TB物理模拟数据集,支持机器学习模型的多样化测试与训练。

The Well: a Large-Scale Collection of Diverse Physics Simulations for Machine Learning

论文配图:The Well: a Large-Scale Collection of Diverse Physics Simulations for Machine Learning
图 1 · 摘自论文原文
  • 整合16个跨领域物理系统模拟数据,覆盖流体、生物、磁流体等复杂动态。
  • 提供统一PyTorch接口,支持模型训练与评估,总数据量达15TB。
  • 适合研究复杂物理系统建模、仿真加速及通用代理模型的开发者。

基于机器学习的代理模型为加速模拟工作流提供了强大工具。然而,现有标准数据集通常仅涵盖少量物理行为类别,难以有效评估新方法性能。为此,我们推出The Well:一个大规模物理模拟数据集合,包含16个数据集,总计15TB,覆盖生物系统、流体动力学、声学散射,以及星系外流体和超新星爆发的磁流体模拟等多样化领域。这些数据可独立使用或作为综合性基准套件。为便于使用,我们提供了统一的PyTorch接口,用于模型训练与评估。通过引入基线模型,展示了The Well所带来复杂动态带来的新挑战。代码与数据已开源(https://github.com/PolymathicAI/the_well)。

原文摘要 · Abstract (English)

Machine learning based surrogate models offer researchers powerful tools for accelerating simulation-based workflows. However, as standard datasets in this space often cover small classes of physical behavior, it can be difficult to evaluate the efficacy of new approaches. To address this gap, we introduce the Well: a large-scale collection of datasets containing numerical simulations of a wide variety of spatiotemporal physical systems. The Well draws from domain experts and numerical software developers to provide 15TB of data across 16 datasets covering diverse domains such as biological systems, fluid dynamics, acoustic scattering, as well as magneto-hydrodynamic simulations of extra-galactic fluids or supernova explosions. These datasets can be used individually or as part of a broader benchmark suite. To facilitate usage of the Well, we provide a unified PyTorch interface for training and evaluating models. We demonstrate the function of this library by introducing example baselines that highlight the new challenges posed by the complex dynamics of the Well. The code and data is available at https://github.com/PolymathicAI/the_well.

物理模拟大规模数据代理模型机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。