arXiv:2510.03191cs.CV2025-10

用产品量化提升图像生成质量,速度更快分辨率更高。

Product-Quantised Image Representation for High-Quality Image Synthesis

  • 将产品量化嵌入VQGAN框架,优化潜空间表示
  • 重建性能大幅提升,PSNR达37dB,FID等指标降96%
  • 可无缝接入扩散模型,加速生成或翻倍输出分辨率

产品量化(PQ)是一种经典的可扩展向量编码方法,但在高保真图像生成的潜空间表示中应用有限。本文提出PQGAN,一种将PQ整合进VQGAN经典向量量化框架的量化图像自编码器。PQGAN在重建性能上显著优于现有先进方法,包括量化与连续对应方法。我们实现了37dB的PSNR,远超此前27dB的水平,并使FID、LPIPS和CMMD得分最高降低96%。成功关键在于对码本大小、嵌入维度与子空间分解之间交互关系的深入分析,其中向量量化和标量量化均为特例。我们发现,当扩展嵌入维度时,VQ与PQ的表现呈相反趋势。此外,该分析揭示了PQ的性能规律,有助于指导超参数最优选择。最后,我们证明PQGAN可无缝集成至预训练扩散模型中,实现更快速、更高效的生成,或在不增加成本的前提下将输出分辨率翻倍,表明PQ是图像合成中离散潜变量表示的强大扩展。

原文摘要 · Abstract (English)

Product quantisation (PQ) is a classical method for scalable vector encoding, yet it has seen limited usage for latent representations in high-fidelity image generation. In this work, we introduce PQGAN, a quantised image autoencoder that integrates PQ into the well-known vector quantisation (VQ) framework of VQGAN. PQGAN achieves a noticeable improvement over state-of-the-art methods in terms of reconstruction performance, including both quantisation methods and their continuous counterparts. We achieve a PSNR score of 37dB, where prior work achieves 27dB, and are able to reduce the FID, LPIPS, and CMMD score by up to 96%. Our key to success is a thorough analysis of the interaction between codebook size, embedding dimensionality, and subspace factorisation, with vector and scalar quantisation as special cases. We obtain novel findings, such that the performance of VQ and PQ behaves in opposite ways when scaling the embedding dimension. Furthermore, our analysis shows performance trends for PQ that help guide optimal hyperparameter selection. Finally, we demonstrate that PQGAN can be seamlessly integrated into pre-trained diffusion models. This enables either a significantly faster and more compute-efficient generation, or a doubling of the output resolution at no additional cost, positioning PQ as a strong extension for discrete latent representation in image synthesis.

图像生成产品量化自编码器扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。