NeurLZ通过在线学习提升科学数据压缩效率,兼顾质量与适应性。
NeurLZ: An Online Neural Learning-Based Method to Enhance Scientific Lossy Compression
- 压缩时在线训练轻量模型,实时适应数据特征变化。
- 前五轮学习即实现89%比特率降低,最高达94%。
- 支持严格或宽松误差约束,适合动态科学数据场景。
大规模科学模拟产生海量数据,存储与读写面临挑战。传统有损压缩难以在压缩比、数据质量和对多样数据特征的适应性间取得平衡。尽管深度学习方法被探索,但普遍依赖大模型和离线训练,限制了对动态数据特性的适应性和计算效率。为此,我们提出NeurLZ,一种融合在线学习、跨场学习和鲁棒误差调控的神经压缩方法。核心创新包括:(1) 压缩时在线学习,采用轻量跳过DNN模型,无需昂贵离线预训练即可适应残差误差;(2) 具备误差补偿能力,可恢复传统压缩器忽略的细节;(3) 支持1×和2×误差调节模式,严格遵守用户输入的1×误差上限或放宽至2×以提升整体质量;(4) 利用科学数据中多场间的相关性进行跨场学习,增强传统方法。在代表性高性能计算数据集(如Nyx、Miranda、Hurricane)上的全面评估表明,NeurLZ在前五个学习周期内实现89%的比特率降低,进一步优化后可达约94%的压缩率,畸变相当,显著优于现有方法,验证了其在提升科学有损压缩方面的高效可扩展性。
原文摘要 · Abstract (English)
Large-scale scientific simulations generate massive datasets, posing challenges for storage and I/O. Traditional lossy compression struggles to advance more in balancing compression ratio, data quality, and adaptability to diverse scientific data features. While deep learning-based solutions have been explored, their common practice of relying on large models and offline training limits adaptability to dynamic data characteristics and computational efficiency. To address these challenges, we propose NeurLZ, a neural method designed to enhance lossy compression by integrating online learning, cross-field learning, and robust error regulation. Key innovations of NeurLZ include: (1) compression-time online neural learning with lightweight skipping DNN models, adapting to residual errors without costly offline pertaining, (2) the error-mitigating capability, recovering fine details from compression errors overlooked by conventional compressors, (3) $1\times$ and $2\times$ error-regulation modes, ensuring strict adherence to $1\times$ user-input error bounds strictly or relaxed 2$\times$ bounds for better overall quality, and (4) cross-field learning leveraging inter-field correlations in scientific data to improve conventional methods. Comprehensive evaluations on representative HPC datasets, e.g., Nyx, Miranda, Hurricane, against state-of-the-art compressors show NeurLZ's effectiveness. During the first five learning epochs, NeurLZ achieves an 89% bit rate reduction, with further optimization yielding up to around 94% reduction at equivalent distortion, significantly outperforming existing methods, demonstrating NeurLZ's superior performance in enhancing scientific lossy compression as a scalable and efficient solution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。