arXiv:2607.15933cs.CV2026-07被引 1

通过匹配特征与码本向量分布,解决向量量化训练不稳和码本坍塌问题。

Distributional Matching for Vector Quantization: A Unified Theoretical and Empirical Framework

论文配图:Distributional Matching for Vector Quantization: A Unified Theoretical and Empirical Framework
图 1 · 摘自论文原文
  • 基于分布对齐思想设计统一的向量量化框架
  • 在多个视觉标记化基准上实现更稳定且鲁棒的性能
  • 适合关注模型训练稳定性与表示效率的研究者

现代视觉表征学习与自回归模型的性能高度依赖于向量量化(VQ),该技术通过可学习的码本对连续特征表示进行离散化。尽管应用广泛,现有VQ方法常因直通估计器引发的梯度不匹配以及码向量利用不足,导致训练不稳定和码本坍塌。本文指出,上述问题根源在于特征向量与码向量分布间的根本性不匹配,造成表示效率低下和信息损失。基于此,我们提出一种分布匹配的向量量化框架,引入理想VQ行为的理论判据,并通过理论分析与实证验证表明,对齐特征与码向量分布可统一缓解训练不稳与码本坍塌。我们基于Wasserstein距离设计目标函数,在温和高斯近似下获得高效闭式解;进一步证明基于最大均值差异(MMD)的非参数方法亦能达到相当效果。在多个视觉标记化基准上的大量实验验证了所提方法的有效性与鲁棒性。

原文摘要 · Abstract (English)

The effectiveness of modern visual representation learning and autoregressive models critically depends on vector quantization (VQ), which discretizes continuous feature representations using a learnable codebook. Despite its widespread use, existing VQ methods often suffer from training instability and codebook collapse, arising from gradient mismatch induced by the straight-through estimator and the under-utilization of code vectors. In this work, we show that both issues can be traced to a fundamental mismatch between the distributions of feature vectors and code vectors, leading to inefficient representation and information loss. Building on this observation, we propose a distributional matching framework for vector quantization. We introduce principled criteria for desirable VQ behavior and demonstrate through theoretical analysis and empirical evaluation that aligning feature and code vector distributions provides a unifying mechanism for mitigating training instability and codebook collapse. We instantiate this framework using a Wasserstein-based objective with an efficient closed-form under a mild Gaussian approximation, and further show that a nonparametric alternative based on maximum mean discrepancy yields comparable performance. Extensive experiments on visual tokenization benchmarks support the effectiveness and robustness of the proposed approach.

向量量化分布对齐训练稳定码本坍塌

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。