用新模型和并行技术,让超算在万亿级网格上精准模拟小尺度湍流。
Pixel-Resolved Long-Context Learning for Turbulence at Exascale: Resolving Small-scale Eddies Toward the Viscous Limit
- 设计分层湍流变换器与环形并行策略,将序列长度从十亿级压缩至百万级。
- 在32768块GPU上实现1.1 EFLOPS算力,缩放效率达94%。
- 首次在AI模型中捕捉到接近粘性极限的小尺度涡旋,适合高精度物理仿真研究者。
湍流在气动、聚变、燃烧等多物理场应用中至关重要。准确捕捉其多尺度特性对可靠预测多物理相互作用至关重要,但即便在百亿亿次超算和先进深度学习模型下仍是一大挑战。需以十亿至万亿级网格点表示的超高分辨率数据,使基于视觉变换器等架构的模型面临难以承受的计算成本。为此,我们提出一种多尺度分层湍流变换器,将序列长度从十亿量级降至百万量级,并引入新型环形序列并行方法(RingX),实现可扩展的长上下文学习。我们在前沿超算上开展缩放与科学计算实验。该方法在32,768块AMD GPU上达到1.1 EFLOPS算力,缩放效率为94%。据我们所知,这是首个能捕捉接近耗散范围小尺度涡旋的湍流人工智能模型。
原文摘要 · Abstract (English)
Turbulence plays a crucial role in multiphysics applications, including aerodynamics, fusion, and combustion. Accurately capturing turbulence's multiscale characteristics is essential for reliable predictions of multiphysics interactions, but remains a grand challenge even for exascale supercomputers and advanced deep learning models. The extreme-resolution data required to represent turbulence, ranging from billions to trillions of grid points, pose prohibitive computational costs for models based on architectures like vision transformers. To address this challenge, we introduce a multiscale hierarchical Turbulence Transformer that reduces sequence length from billions to a few millions and a novel RingX sequence parallelism approach that enables scalable long-context learning. We perform scaling and science runs on the Frontier supercomputer. Our approach demonstrates excellent performance up to 1.1 EFLOPS on 32,768 AMD GPUs, with a scaling efficiency of 94%. To our knowledge, this is the first AI model for turbulence that can capture small-scale eddies down to the dissipative range.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。