统一图像生成与编辑,提升模型通用性与效率
DreamOmni: Unified Image Generation and Editing
- 设计统一框架,整合文本生成与多种编辑任务
- 通过贴纸式合成数据,高效构建高质量编辑训练集
- 联合训练提升生成与编辑性能,适合多任务视觉应用
当前大型语言模型的成功表明,统一的多任务方法能显著提升模型可用性、简化部署并带来跨任务协同效益。然而,在计算机视觉领域,尽管文本到图像(T2I)模型通过规模扩展大幅提升了生成质量,其架构设计最初并未考虑与下游任务(如各类编辑)的融合。为此,我们提出 DreamOmni,一个统一的图像生成与编辑模型。首先分析现有框架与下游任务需求,提出集成 T2I 与多种编辑任务的统一框架。其次,针对指令式与拖拽式编辑数据的高效生成难题,我们设计了一种基于贴纸元素的合成数据流水线,可高效生成准确、高质量的标注数据,实现编辑数据的规模化。在训练阶段,DreamOmni 联合训练生成与编辑任务:生成训练增强模型对特定概念的理解并提升生成质量,编辑训练帮助模型掌握编辑细节。二者协同显著提升编辑表现。大量实验验证了 DreamOmni 的有效性,代码与模型将开源。
原文摘要 · Abstract (English)
Currently, the success of large language models (LLMs) illustrates that a unified multitasking approach can significantly enhance model usability, streamline deployment, and foster synergistic benefits across different tasks. However, in computer vision, while text-to-image (T2I) models have significantly improved generation quality through scaling up, their framework design did not initially consider how to unify with downstream tasks, such as various types of editing. To address this, we introduce DreamOmni, a unified model for image generation and editing. We begin by analyzing existing frameworks and the requirements of downstream tasks, proposing a unified framework that integrates both T2I models and various editing tasks. Furthermore, another key challenge is the efficient creation of high-quality editing data, particularly for instruction-based and drag-based editing. To this end, we develop a synthetic data pipeline using sticker-like elements to synthesize accurate, high-quality datasets efficiently, which enables editing data scaling up for unified model training. For training, DreamOmni jointly trains T2I generation and downstream tasks. T2I training enhances the model's understanding of specific concepts and improves generation quality, while editing training helps the model grasp the nuances of the editing task. This collaboration significantly boosts editing performance. Extensive experiments confirm the effectiveness of DreamOmni. The code and model will be released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。