arXiv:2508.03872cs.LGcs.AI2025-08被引 4

用智能采样大幅减少湍流数据量,提升模型训练效率与精度。

Intelligent Sampling of Extreme-Scale Turbulence Datasets for Accurate and Efficient Spatiotemporal Model Training

  • 基于最大熵原理的智能采样方法,自动筛选关键数据点。
  • 在前沿超算上验证,能耗降低最高达38倍,精度反而提升。
  • 适合大规模物理模拟数据训练,尤其关注能效与精度平衡的研究者。

随着摩尔定律和丹纳德缩放的终结,高效训练亟需重新思考数据规模。我们能否通过智能子采样,用更少的数据训练出更好的模型?为此,我们提出SICKLE——一种用于高效学习的稀疏智能整理框架,包含新颖的最大熵(MaxEnt)采样方法、可扩展训练及能耗基准测试。我们在大型直接数值模拟(DNS)湍流数据集上对比了MaxEnt、随机采样与相空间采样。在前沿超算上对SICKLE进行大规模评估,结果表明,将子采样作为预处理步骤,在多数情况下可提高模型精度并显著降低能耗,最高实现38倍的能耗下降。

原文摘要 · Abstract (English)

With the end of Moore's law and Dennard scaling, efficient training increasingly requires rethinking data volume. Can we train better models with significantly less data via intelligent subsampling? To explore this, we develop SICKLE, a sparse intelligent curation framework for efficient learning, featuring a novel maximum entropy (MaxEnt) sampling approach, scalable training, and energy benchmarking. We compare MaxEnt with random and phase-space sampling on large direct numerical simulation (DNS) datasets of turbulence. Evaluating SICKLE at scale on Frontier, we show that subsampling as a preprocessing step can, in many cases, improve model accuracy and substantially lower energy consumption, with observed reductions of up to 38x.

湍流模拟智能采样能效优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。