ACE++让指令驱动的图像生成与编辑更精准,支持多种任务且训练高效。
ACE++: Instruction-Based Image Creation and Editing via Context-Aware Content Filling
- 基于上下文感知内容填充机制,统一处理各类图像生成与编辑任务。
- 两阶段训练显著降低微调大模型的成本,支持快速适配新任务。
- 提供全量与轻量微调模型,兼顾通用性与垂直场景应用。
我们提出ACE++,一种基于指令的扩散框架,可处理多种图像生成与编辑任务。受FLUX.1-Fill-dev提出的修补任务输入格式启发,改进了ACE中的长上下文条件单元(LCU),并将此输入范式拓展至任意编辑与生成任务。为充分利用图像生成先验知识,设计了两阶段训练方案,以最小化对FLUX.1-dev等强大文本到图像扩散模型的微调成本。第一阶段使用来自文本到图像模型的0参考任务数据进行预训练,许多社区模型如FLUX.1-Fill-dev即符合此范式,可作为初始化加速训练。第二阶段在全部ACE定义的任务上对模型进行微调,以支持通用指令。为促进ACE++在多场景下的广泛应用,提供了涵盖全量微调与轻量微调的完整模型集,兼顾通用性与垂直领域适用性。定性分析显示,ACE++在图像质量与提示遵循能力方面表现优异。代码与模型将发布于项目页:https://ali-vilab.github.io/ACE_plus_page/。
原文摘要 · Abstract (English)
We report ACE++, an instruction-based diffusion framework that tackles various image generation and editing tasks. Inspired by the input format for the inpainting task proposed by FLUX.1-Fill-dev, we improve the Long-context Condition Unit (LCU) introduced in ACE and extend this input paradigm to any editing and generation tasks. To take full advantage of image generative priors, we develop a two-stage training scheme to minimize the efforts of finetuning powerful text-to-image diffusion models like FLUX.1-dev. In the first stage, we pre-train the model using task data with the 0-ref tasks from the text-to-image model. There are many models in the community based on the post-training of text-to-image foundational models that meet this training paradigm of the first stage. For example, FLUX.1-Fill-dev deals primarily with painting tasks and can be used as an initialization to accelerate the training process. In the second stage, we finetune the above model to support the general instructions using all tasks defined in ACE. To promote the widespread application of ACE++ in different scenarios, we provide a comprehensive set of models that cover both full finetuning and lightweight finetuning, while considering general applicability and applicability in vertical scenarios. The qualitative analysis showcases the superiority of ACE++ in terms of generating image quality and prompt following ability. Code and models will be available on the project page: https://ali-vilab. github.io/ACE_plus_page/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。