用深度代理模型快速预测压缩后数据质量,大幅降低计算开销。
DeepCQ: General-Purpose Deep-Surrogate Framework for Lossy Compression Quality Prediction
- 构建通用代理模型,适配多种压缩器和数据集。
- 两阶段设计分离特征提取与预测,训练快、推理轻。
- 专家混合机制提升时序数据预测鲁棒性,误差普遍低于10%。
随着现代科学模拟与仪器生成的数据量持续增长,有界误差的有损压缩技术已成为科学数据管理与分析的关键。然而,压缩后的数据质量评估因指标计算密集而成本高昂。本文提出一种通用的深度代理框架 DeepCQ,用于有损压缩质量预测。其核心贡献包括:1)开发一个可泛化至不同有界误差压缩算法、质量度量及输入数据集的代理模型;2)采用新颖的两阶段设计,将计算昂贵的特征提取与轻量级质量预测解耦,实现高效训练与模块化推理;3)通过专家混合架构优化对时变数据的预测性能,显著提升在不同仿真时间步间差异较大时的预测鲁棒性。我们在四个真实科学应用中验证了 DeepCQ 的有效性,结果表明其预测误差普遍低于10%,显著优于现有方法。该框架使科学用户可根据自身数据质量需求做出明智压缩决策,显著降低科学数据分析中的输入输出与计算开销。
原文摘要 · Abstract (English)
Error-bounded lossy compression techniques have become vital for scientific data management and analytics, given the ever-increasing volume of data generated by modern scientific simulations and instruments. Nevertheless, assessing data quality post-compression remains computationally expensive due to the intensive nature of metric calculations. In this work, we present a general-purpose deep-surrogate framework for lossy compression quality prediction (DeepCQ), with the following key contributions: 1) We develop a surrogate model for compression quality prediction that is generalizable to different error-bounded lossy compressors, quality metrics, and input datasets; 2) We adopt a novel two-stage design that decouples the computationally expensive feature-extraction stage from the light-weight metrics prediction, enabling efficient training and modular inference; 3) We optimize the model performance on time-evolving data using a mixture-of-experts design. Such a design enhances the robustness when predicting across simulation timesteps, especially when the training and test data exhibit significant variation. We validate the effectiveness of DeepCQ on four real-world scientific applications. Our results highlight the framework's exceptional predictive accuracy, with prediction errors generally under 10\% across most settings, significantly outperforming existing methods. Our framework empowers scientific users to make informed decisions about data compression based on their preferred data quality, thereby significantly reducing I/O and computational overhead in scientific data analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。