arXiv:2501.18593cs.CVcs.AI2025-01被引 28

用扩散模型统一训练图像分词器,更简单高效。

Diffusion Autoencoders are Scalable Image Tokenizers

  • 仅用扩散L2损失训练,无需复杂组合损失。
  • 在重建和生成任务中表现优于或持平现有方法。
  • 适合追求简洁、自监督图像建模的研究者。

将图像转化为紧凑视觉表示是构建高效高质量图像生成模型的关键步骤。我们提出一种简单的扩散分词器(DiTo),通过学习紧凑的视觉表示来支持图像生成。核心洞察是:单一的扩散L2损失即可用于训练可扩展的图像分词器。由于扩散模型已在图像生成中广泛使用,这一思路极大简化了分词器的训练流程。相比之下,当前最先进的分词器依赖于经验性设计的损失组合与启发式策略,需复杂训练方案,且依赖预训练监督模型并精细平衡多目标损失。我们通过设计选择与理论分析,使DiTo能够规模化学习具有竞争力的图像表示。实验表明,DiTo在图像重建与下游生成任务中表现优异,是现有监督式分词器的更简单、可扩展且自监督的替代方案。

原文摘要 · Abstract (English)

Tokenizing images into compact visual representations is a key step in learning efficient and high-quality image generative models. We present a simple diffusion tokenizer (DiTo) that learns compact visual representations for image generation models. Our key insight is that a single learning objective, diffusion L2 loss, can be used for training scalable image tokenizers. Since diffusion is already widely used for image generation, our insight greatly simplifies training such tokenizers. In contrast, current state-of-the-art tokenizers rely on an empirically found combination of heuristics and losses, thus requiring a complex training recipe that relies on non-trivially balancing different losses and pretrained supervised models. We show design decisions, along with theoretical grounding, that enable us to scale DiTo for learning competitive image representations. Our results show that DiTo is a simpler, scalable, and self-supervised alternative to the current state-of-the-art image tokenizer which is supervised. DiTo achieves competitive or better quality than state-of-the-art in image reconstruction and downstream image generation tasks.

图像分词扩散模型自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。