提出自适应图像分块方法,让重要区域多分块,不重要的少分块。
Mutual Enhancement Between Global Tokens and Patch Tokens: From Theory to Practice

- 基于信息熵设计动态分块策略,区分图像不同区域的重要性。
- 引入全局令牌与动态过滤算法,减少冗余并提升重建精度。
- 比现有方法快8.7倍,图像质量提升1.3倍,适合长序列图像处理。
准确高效的离散化图像标记对长图像序列处理至关重要。然而,现有方法固定速率压缩所有内容,忽视图像中信息密度的差异,导致冗余或信息丢失。受信息熵启发,我们提出 TaTok——一种理论驱动的自适应图像标记框架。严格分析发现:仅用局部标记重建图像存在信息不足,而局部标记间存在冗余。为此,我们引入全局标记以建模局部标记间的互信息,并设计基于累积条件熵的动态标记过滤(DTF)算法消除冗余。实验验证,TaTok 达到顶尖性能,实现 1.3 倍 gFID 提升和 8.7 倍推理速度加速。通过按信息丰富度分配标记,该方法在更小压缩率下仍保持高精度,为未来研究提供重要启示。
原文摘要 · Abstract (English)
Accurate and effective discrete image tokenization is crucial for long image sequence processing. However, current methods rigidly compress all content at a fixed rate, ignoring the variable information density of images and leading to either redundancy or information loss. Inspired by information entropy, we propose TaTok, a Theoretically grounded adaptive image Tokenization framework. We rigorously identify two key drawbacks in existing methods: information insufficiency when reconstructing images with patch tokens alone, and information redundancy among patch tokens. To address these, we introduce global tokens that model mutual information across patch tokens, and a Dynamic Token Filtering (DTF) algorithm based on cumulative conditional entropy to eliminate redundancy. Experiments confirm TaTok's state-of-the-art performance, delivering a 1.3x gFID improvement and 8.7x inference speedup. By allocating tokens according to information richness, TaTok enables more compressed yet accurate image tokenization, offering valuable insights for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。