arXiv:2603.29634cs.CVcs.AI2026-03被引 1

用遮蔽与语义引导实现高效图像生成的紧凑表征

MacTok: Robust Continuous Tokenization for Image Generation

  • 通过随机遮蔽和语义遮蔽双重机制增强潜空间学习
  • 仅用64或128个令牌即达ImageNet上1.44~1.52的gFID
  • 适合追求高保真且低资源消耗图像生成的研究者

连续图像分词器可提升视觉生成效率,基于变分框架的方法通过KL正则化学习平滑、结构化的潜在表示。然而,在使用较少分词时常出现后验坍缩,导致编码器无法将有效特征压缩至潜空间。为此,我们提出MacTok——一种掩码增强的一维连续分词器,结合图像掩码与表征对齐,防止坍缩的同时学习紧凑鲁棒的表示。MacTok采用随机掩码正则化潜空间学习,并利用DINO引导的语义掩码强调图像中关键区域,迫使模型从不完整视觉证据中编码鲁棒语义。结合全局与局部表征对齐,其在高度压缩的一维潜空间中保留丰富判别信息,仅需64或128个分词。在ImageNet上,以SiT-XL模型实现256×256分辨率下1.44的竞争力gFID,512×512下达1.52的最先进水平,同时减少高达64倍的分词用量。结果表明,掩码与语义引导协同有效防止后验坍缩,实现高效高保真分词。

原文摘要 · Abstract (English)

Continuous image tokenizers enable efficient visual generation, and those based on variational frameworks can learn smooth, structured latent representations through KL regularization. Yet this often leads to posterior collapse when using fewer tokens, where the encoder fails to encode informative features into the compressed latent space. To address this, we introduce \textbf{MacTok}, a \textbf{M}asked \textbf{A}ugmenting 1D \textbf{C}ontinuous \textbf{Tok}enizer that leverages image masking and representation alignment to prevent collapse while learning compact and robust representations. MacTok applies both random masking to regularize latent learning and DINO-guided semantic masking to emphasize informative regions in images, forcing the model to encode robust semantics from incomplete visual evidence. Combined with global and local representation alignment, MacTok preserves rich discriminative information in a highly compressed 1D latent space, requiring only 64 or 128 tokens. On ImageNet, MacTok achieves a competitive gFID of 1.44 at 256$\times$256 and a state-of-the-art 1.52 at 512$\times$512 with SiT-XL, while reducing token usage by up to 64$\times$. These results confirm that masking and semantic guidance together prevent posterior collapse and achieve efficient, high-fidelity tokenization.

图像生成连续分词潜空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。