arXiv:2602.05790cs.ITcs.LG2026-02被引 1

证明了通用码书在向量量化中最多损失0.11比特信息。

Price of metric universality in vector quantization is at most 0.11 bit

  • 提出一种不依赖输入数据统计的通用码书设计思路。
  • 在高斯权矩阵下,通用码书性能仅比最优码书低0.11比特/维。
  • 适用于低精度模型存储,为高效部署大模型提供理论支持。

现代大语言模型中快速计算矩阵乘积 $W^ op X$ 是核心任务。为提升部署效率,常采用低精度近似 $\\( \\widehat W$ 替代真实权重 $W$(即权重仅量化)。信息论表明,最优量化方案需根据 $X$ 的二阶统计特性调整向量量化码书,且与 $X$ 的PCA方向对齐(称作“水填充分配”),但该方法因依赖 $X$ 统计而难以实用。本文证明:存在一个通用码书,可同时近乎最优地适应所有 $X$ 的统计特性,在 $W$ 为高斯分布时,其性能仅比针对 $X$ 优化的水填充码书差 0.11 比特/维。该结果意味着理想低精度存储格式的理论可行性,尽管构造证明是非构造性的。等价地,本结果表明在 $\\mathbb{R}^n$ 中存在一个网,能以几乎最优方式覆盖球面,且对所有希尔伯特范数均成立。

原文摘要 · Abstract (English)

Fast computation of a matrix product $W^\top X$ is a workhorse of modern LLMs. To make their deployment more efficient, a popular approach is that of using a low-precision approximation $\widehat W$ in place of true $W$ (``weight-only quantization''). Information theory demonstrates that an optimal algorithm for reducing precision of $W$ depends on the (second order) statistics of $X$ and requires a careful alignment of vector quantization codebook with PCA directions of $X$ (a process known as ``waterfilling allocation''). Dependence of the codebook on statistics of $X$, however, is highly impractical. This paper proves that there exist a universal codebook that is simultaneously near-optimal for all possible statistics of $X$, in the sense of being at least as good as an $X$-adapted waterfilling codebook with rate reduced by 0.11 bit per dimension in the case when $W$ is Gaussian. Such universal codebook would be an ideal candidate for the low-precision storage format, a topic of active modern research, but alas the existence proof is non-constructive. Equivalently, our result shows existence of a net in $\mathbb{R}^n$ that is a nearly-optimal covering of a sphere simultaneously with respect to all Hilbert norms.

向量量化信息论大模型压缩通用码书

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。