arXiv:2510.22943cs.CV2025-10NeurIPS被引 1

为人脸图像设计可切换的专属码本量化,提升低比特率下的重建质量。

Switchable Token-Specific Codebook Quantization For Face Image Compression

  • 为不同图像类别学习独立码本组,每个令牌分配专属码本。
  • 在0.05 bpp下实现93.51%的人脸识别准确率,显著优于全局码本方法。
  • 设计通用,可嵌入现有码本压缩框架,适合低码率人脸压缩场景。

随着视觉数据量持续增长,高效且无损传输及后续理解已成为现代信息系统的瓶颈。现有基于码本的方法使用全局共享码本对每个令牌进行量化与反量化,通过调整令牌数量或码本大小控制比特率(bpp)。然而,对于富含属性的人脸图像,此类全局码本策略忽略了图像内部类别相关性及令牌间的语义差异,导致低bpp下性能不佳。为此,我们提出可切换的令牌专属码本量化方法,为不同图像类别学习独立码本组,并为每个令牌分配专属码本。通过仅用少量比特记录令牌所属码本组,可降低减小码本组规模带来的损失。这使得在更低总体比特率下仍能支持更多码本,增强表达能力并改善重建效果。由于设计通用,该方法可集成至任意现有码本表示学习框架,在人脸识别数据集上验证有效,于0.05 bpp下实现平均93.51%的重建图像识别准确率。

原文摘要 · Abstract (English)

With the ever-increasing volume of visual data, the efficient and lossless transmission, along with its subsequent interpretation and understanding, has become a critical bottleneck in modern information systems. The emerged codebook-based solution utilize a globally shared codebook to quantize and dequantize each token, controlling the bpp by adjusting the number of tokens or the codebook size. However, for facial images, which are rich in attributes, such global codebook strategies overlook both the category-specific correlations within images and the semantic differences among tokens, resulting in suboptimal performance, especially at low bpp. Motivated by these observations, we propose a Switchable Token-Specific Codebook Quantization for face image compression, which learns distinct codebook groups for different image categories and assigns an independent codebook to each token. By recording the codebook group to which each token belongs with a small number of bits, our method can reduce the loss incurred when decreasing the size of each codebook group. This enables a larger total number of codebooks under a lower overall bpp, thereby enhancing the expressive capability and improving reconstruction performance. Owing to its generalizable design, our method can be integrated into any existing codebook-based representation learning approach and has demonstrated its effectiveness on face recognition datasets, achieving an average accuracy of 93.51% for reconstructed images at 0.05 bpp.

图像压缩码本量化人脸重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。