DreamO统一图像定制框架,支持多条件灵活组合生成。
DreamO: A Unified Framework for Image Customization
- 采用DiT架构统一处理不同类型的输入条件。
- 三阶段训练提升多任务定制能力与生成质量。
- 通过占位符和特征路由实现条件精准控制。
近期大量研究展示了大规模生成模型在图像定制(如身份、主体、风格、背景等)方面的强大能力,但多数方法针对特定任务设计,难以融合多种条件。本文提出DreamO,一个支持多种定制任务且可无缝集成多条件的统一框架。DreamO基于扩散Transformer(DiT)架构,统一处理异构输入;训练时构建涵盖多种定制任务的大规模数据集,并引入特征路由约束,以精确查询参考图像中的相关信息;设计占位符策略,将特定条件绑定至生成结果中的位置,实现条件定位控制;采用三阶段渐进式训练:初期用有限数据完成基础一致性训练,中期全规模训练增强定制能力,末期质量对齐阶段修正低质数据引入的质量偏差。大量实验表明,DreamO能高质量完成多种图像定制任务,并灵活整合多种控制条件。
原文摘要 · Abstract (English)
Recently, extensive research on image customization (e.g., identity, subject, style, background, etc.) demonstrates strong customization capabilities in large-scale generative models. However, most approaches are designed for specific tasks, restricting their generalizability to combine different types of condition. Developing a unified framework for image customization remains an open challenge. In this paper, we present DreamO, an image customization framework designed to support a wide range of tasks while facilitating seamless integration of multiple conditions. Specifically, DreamO utilizes a diffusion transformer (DiT) framework to uniformly process input of different types. During training, we construct a large-scale training dataset that includes various customization tasks, and we introduce a feature routing constraint to facilitate the precise querying of relevant information from reference images. Additionally, we design a placeholder strategy that associates specific placeholders with conditions at particular positions, enabling control over the placement of conditions in the generated results. Moreover, we employ a progressive training strategy consisting of three stages: an initial stage focused on simple tasks with limited data to establish baseline consistency, a full-scale training stage to comprehensively enhance the customization capabilities, and a final quality alignment stage to correct quality biases introduced by low-quality data. Extensive experiments demonstrate that the proposed DreamO can effectively perform various image customization tasks with high quality and flexibly integrate different types of control conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。