用基础模型压缩科学数据,提升效率与质量。
Foundation Model for Lossy Compression of Spatiotemporal Scientific Data
- 结合变分自编码器与超分辨率模块,建模时空依赖关系。
- 微调后压缩比达现有方法4倍,超分辨率提升30%效率。
- 适合大规模科学模拟数据的存储与传输场景。
我们提出一种用于有损科学数据压缩的基础模型(FM),融合变分自编码器(VAE)与超先验结构,并引入超分辨率(SR)模块。VAE利用超先验建模潜在空间依赖性,提升压缩效率;SR模块将低分辨率表示重构为高分辨率输出,改善重建质量。通过交替使用2D与3D卷积,模型高效捕捉科学数据的时空相关性,同时保持低计算开销。实验表明,该模型在未见领域和不同数据形状上均具有良好泛化能力,经领域特定微调后,压缩比最高可达当前最优方法的4倍。相比简单上采样,SR模块使压缩比提升30%。该方法显著降低大规模科学模拟的数据存储与传输成本,同时保障数据完整性与保真度。
原文摘要 · Abstract (English)
We present a foundation model (FM) for lossy scientific data compression, combining a variational autoencoder (VAE) with a hyper-prior structure and a super-resolution (SR) module. The VAE framework uses hyper-priors to model latent space dependencies, enhancing compression efficiency. The SR module refines low-resolution representations into high-resolution outputs, improving reconstruction quality. By alternating between 2D and 3D convolutions, the model efficiently captures spatiotemporal correlations in scientific data while maintaining low computational cost. Experimental results demonstrate that the FM generalizes well to unseen domains and varying data shapes, achieving up to 4 times higher compression ratios than state-of-the-art methods after domain-specific fine-tuning. The SR module improves compression ratio by 30 percent compared to simple upsampling techniques. This approach significantly reduces storage and transmission costs for large-scale scientific simulations while preserving data integrity and fidelity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。