arXiv:2412.07836gr-qcastro-ph.IM2024-12被引 1

用机器学习加速相对论流体模拟中的状态转换,速度提升400倍。

Machine learning-driven conservative-to-primitive conversion in hybrid piecewise polytropic and tabulated equations of state

  • 用神经网络替代传统数值求根法,实现快速状态变量转换。
  • 在百万级数据上比单线程CPU快400倍,误差极低。
  • 适合需要高速模拟的天体物理和高能物理研究者。

我们提出一种新型机器学习方法,用于加速混合分段多项式与查表型物态方程中保守量到原始量的反演。传统数值求根方法计算成本高,尤其在大规模相对论流体模拟中。为此,我们采用前馈神经网络(NNC2PS 和 NNC2PL),在 PyTorch 中训练,并通过 NVIDIA TensorRT 优化以实现 GPU 推理加速。NNC2PS 模型的 $ L_1 $ 与 $ L_ ty $ 误差分别为 $ 4.54 \times 10^{-7} $ 与 $ 3.44 \times 10^{-6} $,NNC2PL 误差更低。使用混合精度部署的 TensorRT 引擎使推理速度相比传统单线程 CPU 实现约 400 倍加速(100万点数据)。在 Delta 超算机上,理想并行下处理 800 万点数据时,TensorRT 相比最优并行数值方法预测可提速 25 倍。该方法还表现出随数据量增长的次线性扩展特性。我们开源了相关科学软件,支持进一步验证与拓展。本工作表明,结合机器学习、GPU 优化与模型量化,可显著加速相对论流体模拟中的保守-原始量转换。

原文摘要 · Abstract (English)

We present a novel machine learning (ML) method to accelerate conservative-to-primitive inversion, focusing on hybrid piecewise polytropic and tabulated equations of state. Traditional root-finding techniques are computationally expensive, particularly for large-scale relativistic hydrodynamics simulations. To address this, we employ feedforward neural networks (NNC2PS and NNC2PL), trained in PyTorch and optimized for GPU inference using NVIDIA TensorRT, achieving significant speedups with minimal accuracy loss. The NNC2PS model achieves $ L_1 $ and $ L_\infty $ errors of $ 4.54 \times 10^{-7} $ and $ 3.44 \times 10^{-6} $, respectively, while the NNC2PL model exhibits even lower error values. TensorRT optimization with mixed-precision deployment substantially accelerates performance compared to traditional root-finding methods. Specifically, the mixed-precision TensorRT engine for NNC2PS achieves inference speeds approximately 400 times faster than a traditional single-threaded CPU implementation for a dataset size of 1,000,000 points. Ideal parallelization across an entire compute node in the Delta supercomputer (Dual AMD 64 core 2.45 GHz Milan processors; and 8 NVIDIA A100 GPUs with 40 GB HBM2 RAM and NVLink) predicts a 25-fold speedup for TensorRT over an optimally-parallelized numerical method when processing 8 million data points. Moreover, the ML method exhibits sub-linear scaling with increasing dataset sizes. We release the scientific software developed, enabling further validation and extension of our findings. This work underscores the potential of ML, combined with GPU optimization and model quantization, to accelerate conservative-to-primitive inversion in relativistic hydrodynamics simulations.

机器学习流体模拟相对论加速

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。