arXiv:2506.15078cs.CVcs.LG2025-06被引 4

用分布对齐提升向量量化,解决训练不稳和代码本坍塌问题。

Enhancing Vector Quantization with Distributional Matching: A Theoretical and Empirical Study

  • 引入Wasserstein距离对齐特征与码本分布
  • 实现接近100%的码本利用率,显著降低量化误差
  • 适合关注量化稳定性与效率的研究者

自回归模型的成功在很大程度上依赖于向量量化技术,该技术通过将连续特征映射到可学习码本中最近的码向量来实现离散化。现有向量量化方法存在两个关键问题:训练不稳定和码本坍塌。训练不稳源于直通估计器带来的梯度偏差,尤其在量化误差较大时;码本坍塌则表现为训练过程中仅使用少量码向量。深入分析表明,这些问题主要由特征分布与码向量分布不匹配引起,导致码向量代表性不足,并在压缩过程中造成显著信息损失。为此,本文采用Wasserstein距离对齐两者分布,实现了接近100%的码本利用率,并显著降低量化误差。实证与理论分析均验证了所提方法的有效性。

原文摘要 · Abstract (English)

The success of autoregressive models largely depends on the effectiveness of vector quantization, a technique that discretizes continuous features by mapping them to the nearest code vectors within a learnable codebook. Two critical issues in existing vector quantization methods are training instability and codebook collapse. Training instability arises from the gradient discrepancy introduced by the straight-through estimator, especially in the presence of significant quantization errors, while codebook collapse occurs when only a small subset of code vectors are utilized during training. A closer examination of these issues reveals that they are primarily driven by a mismatch between the distributions of the features and code vectors, leading to unrepresentative code vectors and significant data information loss during compression. To address this, we employ the Wasserstein distance to align these two distributions, achieving near 100\% codebook utilization and significantly reducing the quantization error. Both empirical and theoretical analyses validate the effectiveness of the proposed approach.

向量量化分布对齐码本坍塌

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。