arXiv:2511.01830cs.LGphysics.flu-dyn2025-11

研究科学计算中高低精度数据如何影响神经代理模型性能,给出高效数据生成策略。

Towards Multi-Fidelity Scaling Laws of Neural Surrogates in CFD

  • 用低/高精度模拟数据构建混合数据集,分解算力预算与数据构成
  • 发现算力增加时模型性能提升有特定规律,存在最优数据精度配比
  • 为科学机器学习中的低成本数据生成提供实证指导,适合领域研究者

缩放定律描述了模型性能随数据量、参数量和算力的增长规律。尽管语言和视觉领域的大规模数据可低成本获取,科学机器学习常受限于数值模拟生成训练数据的高昂成本。通过调整建模假设与近似,仿真精度可与计算成本相互权衡,这是其他领域所不具备的特性。本文利用低精度与高精度雷诺平均纳维-斯托克斯(RANS)模拟,研究神经代理模型中数据精度与成本的权衡关系。重新表述经典缩放定律,将数据集轴分解为算力预算与数据构成。实验揭示了算力-性能的缩放行为,并在给定数据配置下表现出预算依赖的最优精度组合。这些发现首次提供了多精度神经代理数据集的经验缩放定律,为科学机器学习中的算力高效数据生成提供了实践参考。

原文摘要 · Abstract (English)

Scaling laws describe how model performance grows with data, parameters and compute. While large datasets can usually be collected at relatively low cost in domains such as language or vision, scientific machine learning is often limited by the high expense of generating training data through numerical simulations. However, by adjusting modeling assumptions and approximations, simulation fidelity can be traded for computational cost, an aspect absent in other domains. We investigate this trade-off between data fidelity and cost in neural surrogates using low- and high-fidelity Reynolds-Averaged Navier-Stokes (RANS) simulations. Reformulating classical scaling laws, we decompose the dataset axis into compute budget and dataset composition. Our experiments reveal compute-performance scaling behavior and exhibit budget-dependent optimal fidelity mixes for the given dataset configuration. These findings provide the first study of empirical scaling laws for multi-fidelity neural surrogate datasets and offer practical considerations for compute-efficient dataset generation in scientific machine learning.

科学计算神经代理多精度建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。