arXiv:2504.13162cs.CV2025-04被引 7

用自回归模型实现高保真个性化图像生成,效果媲美主流扩散模型。

Personalized Text-to-Image Generation with Auto-Regressive Models

  • 分两阶段优化文本嵌入与Transformer层,提升模型对特定主体的刻画能力。
  • 在主体一致性和提示遵循度上达到与扩散模型相当的水平。
  • 适合关注自回归模型潜力、追求统一模态架构的研究者。

个性化图像生成已成为文本到图像生成中的关键应用,能够生成包含特定主体的多样化场景图像。尽管扩散模型在此领域占据主导地位,但具有统一文本与图像建模架构的自回归模型在个性化生成方面仍鲜被探索。本文研究了优化自回归模型以实现个性化图像合成的潜力,利用其固有的多模态能力完成该任务。我们提出一种两阶段训练策略,结合文本嵌入优化与Transformer层微调。实验表明,该方法在自回归模型上实现了与领先扩散基个性化方法相当的主体保真度和提示遵循能力。结果证明自回归模型在个性化图像生成中的有效性,为该领域未来研究提供了新方向。

原文摘要 · Abstract (English)

Personalized image synthesis has emerged as a pivotal application in text-to-image generation, enabling the creation of images featuring specific subjects in diverse contexts. While diffusion models have dominated this domain, auto-regressive models, with their unified architecture for text and image modeling, remain underexplored for personalized image generation. This paper investigates the potential of optimizing auto-regressive models for personalized image synthesis, leveraging their inherent multimodal capabilities to perform this task. We propose a two-stage training strategy that combines optimization of text embeddings and fine-tuning of transformer layers. Our experiments on the auto-regressive model demonstrate that this method achieves comparable subject fidelity and prompt following to the leading diffusion-based personalization methods. The results highlight the effectiveness of auto-regressive models in personalized image generation, offering a new direction for future research in this area.

自回归模型个性化生成图像合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。