针对科学数据多变量相关性,提出可控制误差的压缩方法
Multivariate Scientific Data Compression with Learned Cross-Variable Latent Decorrelation and Autoregressive Entropy Modeling

- 用可学习正交变换分离变量间依赖关系
- 结合自回归模型捕捉残余空间结构,提升压缩率
- 适合气候、燃烧等多变量科学数据压缩任务
科学模拟产生具有异质统计特性和依赖关系的物理场集合,而现有学习型压缩器通常独立编码各场或依赖共享编码器但未显式建模潜在空间中的结构。本文提出CAESAR-LDAR,一种误差可控的多变量学习压缩器,基于共享的CAESAR-V主干网络,引入两种互补机制:可训练的正交变换重新组织对齐潜变量通道间的依赖关系;因果自回归分层先验捕捉变换后残留的局部空间结构。通过矩阵指数参数化保持正交性,实现精确可逆且无需额外惩罚项。所有变体统一应用残差修正阶段以满足指定重建容差。在燃烧、气候和湍流数据上的实验表明,两种机制在不同场景下均有效:当非线性编码器后仍存在显著线性跨通道依赖时,潜变量去相关作用更明显;当剩余结构主要为局部或空间特征时,自回归建模依然有效。二者结合在所测数据集上达到最优或接近最优的率失真性能。全局变换计算开销极小,而自回归编码带来较大吞吐量代价。结果表明,多变量科学压缩的实用设计原则是:若潜空间中可观测到全局跨通道依赖,则应加以利用;否则以局部概率上下文作为通用补充机制。
原文摘要 · Abstract (English)
Scientific simulations generate collections of physical fields with heterogeneous statistics and dependencies, yet learned compressors often encode those fields independently or rely on a shared encoder without explicitly modeling the structure that remains in latent space. We present CAESAR-LDAR, an error-controlled multivariate learned compressor that augments a shared CAESAR-V backbone with two complementary mechanisms: a trainable orthogonal transform that reorganizes dependence across aligned latent channels, and a causal autoregressive hierarchical prior that captures local spatial structure left after transformation. Orthogonality is maintained through a matrix-exponential parameterization, making the transform exactly invertible without an additional penalty. A common residual-correction stage is applied uniformly to all variants to enforce the requested reconstruction tolerance. Experiments across combustion, climate, and turbulence data show that the two mechanisms are useful in different regimes. Latent decorrelation helps most when substantial linear cross-channel dependence survives the nonlinear encoder, whereas autoregressive modeling remains effective when the remaining structure is primarily local or spatial. Their combination provides the strongest or near-strongest rate-distortion performance across the evaluated datasets. The global transform adds little computational overhead, while autoregressive coding introduces a larger throughput tradeoff. More broadly, the results suggest a practical design principle for multivariate scientific compression: exploit global cross-channel dependence when it is measurably present in latent space, and use local probabilistic context as a complementary mechanism across a wider range of data regimes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。