arXiv:2606.15959cs.DCcs.AI2026-06

用压缩数据训练生成模型,能大幅省存储且提速,误差可控。

Quantifying the Impact of Lossy Compression on Neural Generative Surrogate Modeling

论文配图:Quantifying the Impact of Lossy Compression on Neural Generative Surrogate Modeling
图 1 · 摘自论文原文
  • 通过分析训练不确定性,量化压缩可容忍的误差范围。
  • 压缩后存储减少23.7~39倍,训练速度提升最高3倍。
  • 适合需要高效训练大规模科学模拟替代模型的研究者。

神经网络常作为科学发现中的生成代理模型,是对复杂数值模拟的可训练近似,能替代耗时的仿真获得快速解。然而高保真代理模型需海量训练数据,带来存储与读写挑战。有损压缩有望缓解此问题,但压缩误差可能以微妙方式影响模型质量,难以量化。本文研究训练数据有损压缩对生成代理模型质量的影响。首先揭示神经网络训练中固有的不确定性——相同配置可产生不同模型。利用这一特性,提出一种方法估算压缩引入误差的容忍度。在两个应用仿真上的评估表明,该方法显著降低内存/存储需求并加速训练,同时保持高质量模型。结果表明,压缩可节省高达23.7倍和39倍的数据存储,对模型质量影响微乎其微;同时减小数据集规模可提升数据加载速度,使训练时间最多缩短3倍。

原文摘要 · Abstract (English)

Neural networks are used as generative surrogate models for scientific discovery, which are trainable approximations of scientific simulations. These models enable users to replace time-consuming numerical simulations with learned alternatives, providing quick solutions. However, high-fidelity generative surrogate models require massive training datasets, which can create storage and I/O challenges. Lossy compression is a promising way to reduce this burden, but compression errors may affect the model quality in subtle ways, making it challenging to quantify their impact. In this work, we examine how lossy compression of training data impacts the quality of generative surrogate models. We begin by characterizing the uncertainty inherent in training neural networks, showing that identical training configurations can produce different models. By exploiting this variability, we propose a method to estimate how much compression-induced error a surrogate model can tolerate without affecting its accuracy. Evaluation of two application simulations demonstrates that our approach significantly reduces memory/storage requirements and speeds up training while producing high-quality surrogate models. These results show that lossy compression saves data storage up to 23.7x and 39x with negligible impact on the quality of the surrogate model. Meanwhile, reducing the size of the training data set also enhances the data loading speed and reduces the training time by up to 3x.

生成模型数据压缩科学计算神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。