arXiv:2608.00053eess.IVcs.CV2026-08

用可训练的张量网络优化图像压缩基,提升压缩效率

Fast Trainable Multilinear Bases for Image Compression

论文配图:Fast Trainable Multilinear Bases for Image Compression
图 1 · 摘自论文原文
  • 用量子多体理论启发的等距张量网络参数化压缩基
  • 在手绘图数据集上比传统JPEG的余弦变换快20%
  • 适合追求高压缩率的图像编码研究者

离散傅里叶变换(DFT)、离散余弦变换(DCT)及其分块变体是当前主流图像与视频编码器的核心。它们的有效性依赖于三个特性:运行时间接近线性(最多含对数因子),完全可逆,且参数极少。本文将这些基推广为等距多线性基,在保持上述三特性的同时,引入少量额外参数(图像尺寸的对数级)。我们提出一种方案,针对特定图像数据集训练更优的变换:使用受量子多体理论启发的等距张量网络参数化基,并通过黎曼优化进行训练。实验表明,训练能持续提升性能,所学基可表示传统的DFT与DCT-IV。在自然照片和线条画数据集上均取得成效。例如,在Quick Draw线条画压缩任务中,最佳训练基比JPEG中使用的分块余弦变换减少20%的数据量。

原文摘要 · Abstract (English)

The Discrete Fourier Transform (DFT), the Discrete Cosine Transform (DCT), and their block-wise variants underpin most deployed image and video codecs. Their effectiveness rests on three properties: their runtime is near-linear (up to a polylogarithmic factor) in the image size, they are exactly invertible, and they carry few to no parameters. In this work, we generalize these bases to isometric multilinear bases, allowing a small number of extra parameters (polylogarithmic in the image size), while preserving all three properties. We develop a scheme to train a better transformation for a given image dataset: we use isometric tensor networks, inspired by quantum many-body theory, to parameterize the basis, and train it with Riemannian optimization. We show that training consistently improves performance, as our parameterized bases can represent the traditional DFT and DCT-IV (a variant of the DCT). Evidence is shown across natural photographs and line drawings. On Quick Draw line-drawing compression, for example, the best trained basis outperforms the block cosine transform used in the JPEG format by $20\%$ in terms of compressed data size.

图像压缩张量网络可训练基黎曼优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。