arXiv:2503.12652cs.CV2025-03ICCV被引 21

一个模型搞定图像生成与编辑,无需多个专用模型。

UniVG: A Generalist Diffusion Model for Unified Image Generation and Editing

  • 用统一条件输入支持多种图像任务,仅需一套参数。
  • 在多个任务上表现超越专用模型,无性能损失。
  • 适合需要多场景图像生成的开发者与研究者。

文本到图像(T2I)扩散模型在生成符合用户提示的视觉逼真图像方面表现优异。在此基础上,现有方法通过微调预训练模型实现特定任务,但需独立模型架构、训练设计和参数集。本文提出UniVG,一种通用扩散模型,仅用一套权重即可支持多样化的图像生成任务。UniVG将多模态输入视为统一条件,适用于从T2I生成、修复、基于指令的编辑、身份保持生成、布局引导生成,到深度估计和指代分割等多种下游应用。通过大规模数据混合与多任务训练的实证研究,我们揭示了训练过程中的关键决策。例如,T2I生成与其他任务(如基于指令的编辑)可共存且无性能折损;辅助任务如深度估计和指代分割能提升图像编辑效果。值得注意的是,该模型在部分任务上甚至优于专用模型,标志着向统一图像生成迈进的重要一步。

原文摘要 · Abstract (English)

Text-to-Image (T2I) diffusion models have shown impressive results in generating visually compelling images following user prompts. Building on this, various methods further fine-tune the pre-trained T2I model for specific tasks. However, this requires separate model architectures, training designs, and multiple parameter sets to handle different tasks. In this paper, we introduce UniVG, a generalist diffusion model capable of supporting a diverse range of image generation tasks with a single set of weights. UniVG treats multi-modal inputs as unified conditions to enable various downstream applications, ranging from T2I generation, inpainting, instruction-based editing, identity-preserving generation, and layout-guided generation, to depth estimation and referring segmentation. Through comprehensive empirical studies on data mixing and multi-task training, we provide detailed insights into the training processes and decisions that inform our final designs. For example, we show that T2I generation and other tasks, such as instruction-based editing, can coexist without performance trade-offs, while auxiliary tasks like depth estimation and referring segmentation enhance image editing. Notably, our model can even outperform some task-specific models on their respective benchmarks, marking a significant step towards a unified image generation model.

扩散模型图像生成通用模型编辑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。