提出新型图像分词器GSQ,实现高效高质图像压缩与可扩展性。
Scaling Image Tokenizers with Grouped Spherical Quantization
- 采用球面初始化与查找正则化,约束码本潜空间在球面上。
- 16倍下采样下重建FID达0.50,优于现有方法且训练更少迭代。
- 适用于追求高质量图像压缩与模型可扩展性的研究者。
视觉分词器因其可扩展性和紧凑性受到广泛关注;但此前工作依赖过时的GAN超参数、存在偏差的比较,且缺乏对缩放行为的全面分析。为此,本文提出分组球面量化(GSQ),通过球面码本初始化和查找正则化,将码本潜空间约束在球面。对图像分词器训练策略的实证分析表明,GSQ-GAN在更少训练迭代中实现优于当前最优方法的重建质量,为缩放研究奠定基础。在此基础上,系统考察了GSQ在潜在维度、码本大小和压缩比上的缩放特性及其对模型性能的影响。研究发现,在高、低空间压缩水平下表现出截然不同的行为,凸显高维潜空间表示的挑战。结果表明,GSQ能将高维潜空间重构为紧凑的低维空间,从而实现高效缩放并提升质量。最终,GSQ-GAN在16倍下采样下达到重建FID(rFID)0.50。
原文摘要 · Abstract (English)
Vision tokenizers have gained a lot of attraction due to their scalability and compactness; previous works depend on old-school GAN-based hyperparameters, biased comparisons, and a lack of comprehensive analysis of the scaling behaviours. To tackle those issues, we introduce Grouped Spherical Quantization (GSQ), featuring spherical codebook initialization and lookup regularization to constrain codebook latent to a spherical surface. Our empirical analysis of image tokenizer training strategies demonstrates that GSQ-GAN achieves superior reconstruction quality over state-of-the-art methods with fewer training iterations, providing a solid foundation for scaling studies. Building on this, we systematically examine the scaling behaviours of GSQ, specifically in latent dimensionality, codebook size, and compression ratios, and their impact on model performance. Our findings reveal distinct behaviours at high and low spatial compression levels, underscoring challenges in representing high-dimensional latent spaces. We show that GSQ can restructure high-dimensional latent into compact, low-dimensional spaces, thus enabling efficient scaling with improved quality. As a result, GSQ-GAN achieves a 16x down-sampling with a reconstruction FID (rFID) of 0.50.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。