arXiv:2606.11096cs.CV2026-06被引 1

通过深度对齐浅层与深层特征,提升图像自编码的细节保真度。

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder

论文配图:IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder
图 1 · 摘自论文原文
  • 联合对齐浅层与深层视觉特征,增强离散表示的完整性
  • ImageNet上rFID达0.61,较前人方法提升0.28
  • 适用于高质量自回归图像生成,性能刷新纪录

基于预训练视觉基础模型(VFMs),表示自编码器(RAEs)近年来成为构建语义丰富潜在空间以支持图像生成的有前景方法。然而,其重建质量常不理想,主要因为深层VFM特征缺乏足够的细粒度视觉细节。这一问题在离散化后尤为严重,低层信息缺失难以恢复。我们观察到,浅层VFM特征保留了更丰富的局部外观与结构细节,可补足现有RAE中深层特征所承载的高层语义。受此互补特性启发,我们提出Ideal——一种深度对齐框架,用于离散表示自编码。通过同时对齐量化标记与浅层及深层VFM特征,Ideal使离散视觉标记兼具视觉保真度与丰富语义。大量实验表明,Ideal实现更优重建性能,在ImageNet上rFID达0.61,优于此前最佳方法0.28;用于自回归图像生成时,进一步获得gFID 1.89,创下自回归图像生成新纪录。

原文摘要 · Abstract (English)

Built on pretrained vision foundation models (VFMs), representation autoencoders (RAEs) have recently emerged as a promising approach for constructing semantically rich latent spaces for image generation. However, their reconstruction quality often remains suboptimal, largely because deep VFM representations do not preserve sufficient fine-grained visual detail. This limitation becomes even more severe after discretization, where missing low-level information is difficult to recover. In fact, we observe that shallow VFM features retain considerably richer local appearance and structural detail, which complements the high-level semantics carried by deep features used in existing RAEs. Motivated by this complementary property, we propose Ideal, an In-depth Alignment framework for discrete representation autoencoding. By jointly aligning quantized tokens with both shallow and deep VFM features, Ideal enables the resulting discrete visual tokens to preserve both visual fidelity and rich semantics. Extensive experiments demonstrate that Ideal yields superior reconstruction performance, achieving 0.61 rFID on ImageNet and outperforming the previous best method by 0.28. When used for autoregressive image generation, Ideal further produces a gFID of 1.89, establishing a new state of the art for autoregressive image generation.

自编码器视觉表征图像生成离散化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。