首次系统揭示向量量化中的表征坍缩问题及其成因
Representation Collapsing Problems in Vector Quantization
- 通过合成与真实数据验证表征坍缩的两种类型
- 发现初始化受限和编码器容量不足是主因
- 提出针对性缓解方案,适合生成模型研究者
向量量化是一种将连续表示离散化为一组离散向量的技术,在大语言模型、扩散模型等生成模型中广泛应用。尽管如此,其在生成模型中的特性与行为仍缺乏深入研究。本文首次系统研究了向量量化中的表征坍缩问题——即码本向量或潜在嵌入收敛至有限子集,丧失区分能力,严重影响模型捕捉多样数据模式的能力。通过合成与真实数据集分析,我们揭示了两类坍缩的严重程度及触发条件:受限初始化导致码本坍缩,编码器容量不足引发嵌入坍缩。基于此,我们提出针对性缓解策略。这是首个全面探讨向量量化表征坍缩问题的研究。
原文摘要 · Abstract (English)
Vector quantization is a technique in machine learning that discretizes continuous representations into a set of discrete vectors. It is widely employed in tokenizing data representations for large language models, diffusion models, and other generative models. Despite its prevalence, the characteristics and behaviors of vector quantization in generative models remain largely underexplored. In this study, we investigate representation collapse in vector quantization - a critical degradation where codebook tokens or latent embeddings lose their discriminative power by converging to a limited subset of values. This collapse fundamentally compromises the model's ability to capture diverse data patterns. By leveraging both synthetic and real datasets, we identify the severity of each type of collapses and triggering conditions. Our analysis reveals that restricted initialization and limited encoder capacity result in tokens collapse and embeddings collapse. Building on these findings, we propose potential solutions aimed at mitigating each collapse. To the best of our knowledge, this is the first comprehensive study examining representation collapsing problems in vector quantization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。