arXiv:2609.02107cs.LGcs.CV2026-09

统一量化评估框架,揭示向量量化最优的内在原因

A Unified Rate-Distortion Perspective on Vector, Product, and Scalar Quantization

论文配图:A Unified Rate-Distortion Perspective on Vector, Product, and Scalar Quantization
图 1 · 摘自论文原文
  • 将量化视为有损压缩,用码本大小和令牌数定义码率
  • 证明最小化失真比提高码本利用率更能保证重建质量
  • 提出公平比较标准,适用于对比不同量化方法

离散视觉标记化主要依赖向量、标量和乘积量化,但缺乏统一的理解框架来把握量化权衡。本文提出一种统一的率失真视角,将量化视为有损压缩:以令牌数量和码本大小表征固定长度编码率,量化误差作为失真。在此框架下,解决三个核心问题:第一,理论与实证表明,最小化失真而非最大化码本利用率才是重建保真度的核心目标,且与STE引起的梯度偏差直接相关;第二,确立量化比较的两个关键公平条件:控制潜在特征统计分布并强制相同编码率;第三,在此条件下,恢复出现代视觉标记化中VQ-PQ-SQ的失真层级关系,并实证显示现代VQ方法达到最低失真。本工作为现代离散视觉标记化提供了基础的率失真重构,解决了量化评估中的模糊性,并在固定码率约束下提供可控的内在有效性分析框架。

原文摘要 · Abstract (English)

Discrete visual tokenization, predominantly driven by vector, scalar, and product quantization, lacks a unified conceptual framework for understanding quantization tradeoffs. In this paper, we propose a unified rate--distortion perspective on modern discrete visual tokenization. By viewing quantization as lossy compression, we characterize the nominal fixed-length coding rate through token count and codebook size, and quantization error as the distortion. Within this framework, we resolve three central questions. First, we theoretically and empirically show that minimizing distortion, rather than maximizing codebook utilization, is the primary intrinsic objective for reconstruction fidelity, with a direct connection to the STE-induced gradient discrepancy. Second, we establish two critical fairness conditions for intrinsic quantization comparison: controlling latent feature statistics and enforcing identical coding rates. Third, under these conditions, we recover the VQ--PQ--SQ distortion hierarchy in modern visual tokenization and show empirically that modern VQ methods achieve the lowest distortion. This work provides a foundational rate--distortion reframing of modern discrete visual tokenization, resolves ambiguities in quantizer evaluation, and provides a controlled framework for isolating intrinsic quantization effectiveness under fixed-rate constraints.

量化率失真视觉标记化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。