arXiv:2411.18616cs.CVcs.AI2024-11CVPR被引 56

用自生成数据让文生图模型实现零样本精准图像定制

Diffusion Self-Distillation for Zero-Shot Customized Image Generation

  • 用预训练模型自动生成图文配对数据集
  • 在无需微调的情况下实现身份保留生成效果超越现有方法
  • 适合需要精细控制图像风格与内容的创作者使用

文生图扩散模型虽效果出色,但艺术家难以实现细粒度控制。例如,在新场景中生成特定实例的图像(即身份保留生成)这一任务,天然适合图文条件生成模型。然而,直接训练此类模型缺乏高质量配对数据。本文提出扩散自蒸馏方法,利用预训练文生图模型自动生成用于图文条件图像生成的数据集。首先借助文生图模型的上下文生成能力,结合视觉语言模型筛选,构建大规模图文配对数据集;随后基于该数据集对模型进行微调,使其从文到图生成转变为图文到图生成。实验表明,该方法在多种身份保留生成任务上优于现有零样本方法,且性能接近每实例微调技术,同时无需测试时优化。

原文摘要 · Abstract (English)

Text-to-image diffusion models produce impressive results but are frustrating tools for artists who desire fine-grained control. For example, a common use case is to create images of a specific instance in novel contexts, i.e., "identity-preserving generation". This setting, along with many other tasks (e.g., relighting), is a natural fit for image+text-conditional generative models. However, there is insufficient high-quality paired data to train such a model directly. We propose Diffusion Self-Distillation, a method for using a pre-trained text-to-image model to generate its own dataset for text-conditioned image-to-image tasks. We first leverage a text-to-image diffusion model's in-context generation ability to create grids of images and curate a large paired dataset with the help of a Visual-Language Model. We then fine-tune the text-to-image model into a text+image-to-image model using the curated paired dataset. We demonstrate that Diffusion Self-Distillation outperforms existing zero-shot methods and is competitive with per-instance tuning techniques on a wide range of identity-preservation generation tasks, without requiring test-time optimization.

图像生成扩散模型零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。