用可调控的视觉区块实现像素级图像元素编辑
BlobCtrl: Taming Controllable Blob for Element-level Image Editing
- 将图像拆解为可独立控制的视觉区块,分离布局与外观
- 支持添加、删除、缩放等操作,效果优于现有方法
- 适合需要精细编辑图像元素的研究者和设计师
随着用户对图像编辑要求的提升,当前基于扩散模型的方法在灵活、细粒度操控特定视觉元素方面面临挑战。本文提出BlobCtrl,一种基于概率区块表示的元素级图像编辑框架。将区块视为视觉基本单元,该方法解耦布局与外观,实现可控的对象级操作。主要贡献包括:(1) 一种上下文感知的双分支扩散模型,通过区块表示显式分离前景与背景处理,解耦布局与外观;(2) 一种自监督的解耦-重建训练范式,结合保持身份的损失函数,并设计策略高效利用区块-图像配对数据。为促进研究,我们构建了用于大规模训练的BlobData数据集和用于系统评估的BlobBench基准。实验表明,BlobCtrl在对象增删、缩放、替换等多种元素级编辑任务中达到领先性能,同时保持高效计算。
原文摘要 · Abstract (English)
As user expectations for image editing continue to rise, the demand for flexible, fine-grained manipulation of specific visual elements presents a challenge for current diffusion-based methods. In this work, we present BlobCtrl, a framework for element-level image editing based on a probabilistic blob-based representation. Treating blobs as visual primitives, BlobCtrl disentangles layout from appearance, affording fine-grained, controllable object-level manipulation. Our key contributions are twofold: (1) an in-context dual-branch diffusion model that separates foreground and background processing, incorporating blob representations to explicitly decouple layout and appearance, and (2) a self-supervised disentangle-then-reconstruct training paradigm with an identity-preserving loss function, along with tailored strategies to efficiently leverage blob-image pairs. To foster further research, we introduce BlobData for large-scale training and BlobBench, a benchmark for systematic evaluation. Experimental results demonstrate that BlobCtrl achieves state-of-the-art performance in a variety of element-level editing tasks, such as object addition, removal, scaling, and replacement, while maintaining computational efficiency. Project Webpage: https://liyaowei-stu.github.io/project/BlobCtrl/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。