arXiv:2409.05357cs.LG2024-09被引 8

用注意力机制提升科学数据压缩率,兼顾精度与效率。

Attention Based Machine Learning Methods for Data Reduction with Guaranteed Error Bounds

  • 分块设计+注意力机制捕捉跨块相关性
  • 在多变量数据上压缩率最高达SZ3的8倍
  • 适合高维科学模拟数据的高效压缩

高能物理、计算流体动力学和气候科学等领域的科学应用产生海量高速数据,其增长速度已超过算力、网络和存储的发展。为应对挑战,需采用数据压缩或缩减技术。这些科学数据具有结构化和分块结构化的多维网格特征,每个网格点对应一个张量。数据缩减方法应充分利用普遍存在的空间与时间相关性。此外,如CFD中涉及上百种物质及其属性的流程张量,缩减方法还需利用张量内元素间的相互关系。本文提出一种基于注意力的分层压缩方法,采用分块压缩框架:引入注意力驱动的超块自编码器以捕捉跨块相关性,再通过分块编码器提取块内特异性信息,并结合基于PCA的后处理步骤,确保每一块的数据误差有界。该方法有效捕获了块内及块间时空与变量间相关性。相较于当前最优的SZ3方法,在多变量S3D数据集上压缩率最高提升8倍;在单变量设置下,对E3SM和XGC数据集分别实现最高3倍和2倍的压缩率提升。

原文摘要 · Abstract (English)

Scientific applications in fields such as high energy physics, computational fluid dynamics, and climate science generate vast amounts of data at high velocities. This exponential growth in data production is surpassing the advancements in computing power, network capabilities, and storage capacities. To address this challenge, data compression or reduction techniques are crucial. These scientific datasets have underlying data structures that consist of structured and block structured multidimensional meshes where each grid point corresponds to a tensor. It is important that data reduction techniques leverage strong spatial and temporal correlations that are ubiquitous in these applications. Additionally, applications such as CFD, process tensors comprising hundred plus species and their attributes at each grid point. Reduction techniques should be able to leverage interrelationships between the elements in each tensor. In this paper, we propose an attention-based hierarchical compression method utilizing a block-wise compression setup. We introduce an attention-based hyper-block autoencoder to capture inter-block correlations, followed by a block-wise encoder to capture block-specific information. A PCA-based post-processing step is employed to guarantee error bounds for each data block. Our method effectively captures both spatiotemporal and inter-variable correlations within and between data blocks. Compared to the state-of-the-art SZ3, our method achieves up to 8 times higher compression ratio on the multi-variable S3D dataset. When evaluated on single-variable setups using the E3SM and XGC datasets, our method still achieves up to 3 times and 2 times higher compression ratio, respectively.

数据压缩注意力机制科学计算张量压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。