arXiv:2503.08685cs.CV2025-03ICCV被引 22

用主成分思想生成可解释的图像令牌序列

"Principal Components" Enable A New Language of Images

  • 将主成分分析结构引入隐空间令牌,保证每步添加新信息
  • 重建性能达顶尖水平,且令牌信息量逐级递减
  • 适合需要可解释性的图像生成与理解任务

我们提出一种新型视觉令牌化框架,将可证明的PCA类结构嵌入隐空间令牌中。现有视觉令牌化器多关注重建保真度,却常忽视隐空间的结构性——这对可解释性和下游任务至关重要。本方法生成一维因果令牌序列,每个后续令牌贡献非重叠信息,且数学上保证解释方差递减,类似主成分分析。该结构确保模型优先提取显著视觉特征,后续令牌提供递减但互补的信息。此外,我们识别并解决了语义-频谱耦合效应,该效应导致高层语义与低层频谱细节在令牌中纠缠,通过扩散解码器得以缓解。实验表明,该方法在重建性能上达到当前最优,且更符合人类视觉系统,具备更好可解释性。基于此令牌序列训练的自回归模型性能媲美最先进方法,同时减少训练与推理所需令牌数。

原文摘要 · Abstract (English)

We introduce a novel visual tokenization framework that embeds a provable PCA-like structure into the latent token space. While existing visual tokenizers primarily optimize for reconstruction fidelity, they often neglect the structural properties of the latent space--a critical factor for both interpretability and downstream tasks. Our method generates a 1D causal token sequence for images, where each successive token contributes non-overlapping information with mathematically guaranteed decreasing explained variance, analogous to principal component analysis. This structural constraint ensures the tokenizer extracts the most salient visual features first, with each subsequent token adding diminishing yet complementary information. Additionally, we identified and resolved a semantic-spectrum coupling effect that causes the unwanted entanglement of high-level semantic content and low-level spectral details in the tokens by leveraging a diffusion decoder. Experiments demonstrate that our approach achieves state-of-the-art reconstruction performance and enables better interpretability to align with the human vision system. Moreover, autoregressive models trained on our token sequences achieve performance comparable to current state-of-the-art methods while requiring fewer tokens for training and inference.

图像令牌化主成分分析可解释性自回归模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。