arXiv:2412.10316cs.CVcs.AI2024-12被引 34

用AI画画时能直接说要求,自动改图还保持细节一致。

BrushEdit: All-In-One Image Inpainting and Editing

  • 结合大语言模型与双分支补图模型,实现自由指令编辑
  • 在7项指标上表现优于现有方法,尤其保留修改区域结构
  • 适合希望直接用自然语言操作图像的设计师和开发者

图像编辑随着扩散模型的发展取得了显著进展,目前主要依赖基于反演和基于指令的方法。然而,现有反演方法因反演噪声具有结构性,在大幅修改(如增删物体)时表现不佳;而基于指令的方法常将用户限制在黑箱操作中,难以直接指定编辑区域和强度。为此,我们提出BrushEdit,一种基于补图的指令引导图像编辑新范式,通过多模态大语言模型(MLLMs)与图像补图模型协同工作,实现自主、用户友好且可交互的自由形式指令编辑。具体而言,我们构建了一个代理协作框架,集成MLLMs与双分支图像补图模型,完成编辑类别分类、主体识别、掩码获取及编辑区域补图。大量实验表明,该框架有效融合了MLLMs与补图模型,在包括掩码区域保真度和编辑效果连贯性在内的七项指标上均取得优异表现。

原文摘要 · Abstract (English)

Image editing has advanced significantly with the development of diffusion models using both inversion-based and instruction-based methods. However, current inversion-based approaches struggle with big modifications (e.g., adding or removing objects) due to the structured nature of inversion noise, which hinders substantial changes. Meanwhile, instruction-based methods often constrain users to black-box operations, limiting direct interaction for specifying editing regions and intensity. To address these limitations, we propose BrushEdit, a novel inpainting-based instruction-guided image editing paradigm, which leverages multimodal large language models (MLLMs) and image inpainting models to enable autonomous, user-friendly, and interactive free-form instruction editing. Specifically, we devise a system enabling free-form instruction editing by integrating MLLMs and a dual-branch image inpainting model in an agent-cooperative framework to perform editing category classification, main object identification, mask acquisition, and editing area inpainting. Extensive experiments show that our framework effectively combines MLLMs and inpainting models, achieving superior performance across seven metrics including mask region preservation and editing effect coherence.

图像编辑多模态扩散模型自然语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。