用粗略遮罩精准替换任意物体,支持500万张多类别图像编辑。
A$^2$-Edit: Precise Reference-Guided Image Editing of Arbitrary Objects and Ambiguous Masks
- 通过动态专家选择的Transformer模块区分不同物体类别。
- 在VITON-HD等数据集上超越现有方法,提升跨类别编辑效果。
- 适合需要灵活编辑复杂场景中任意物体的研究与应用者。
我们提出A^2-Edit,一种统一的图像修复框架,可对任意物体类别进行编辑,仅需粗略掩码即可将目标区域替换为参考物体。为解决现有数据集同质化严重和类别覆盖有限的问题,我们构建了大规模多类别数据集UniEdit-500K,包含8大类、209个细粒度子类,共500,104对图像。丰富的类别多样性对模型提出新挑战,要求其自动学习跨类语义关系与差异。为此,我们引入Mixture of Transformer模块,通过动态专家选择实现不同类别的差异化建模,并借助专家间协作增强跨类别语义迁移与泛化能力。此外,我们提出掩码渐进松弛训练策略(MATS),在训练中逐步放宽掩码精度要求,降低模型对精确掩码的依赖,提升在多样化编辑任务中的鲁棒性。在VITON-HD和AnyInsertion等基准上的大量实验表明,A^2-Edit在所有指标上持续优于现有方法,为任意物体编辑提供了高效新方案。
原文摘要 · Abstract (English)
We propose A^2-Edit, a unified inpainting framework for arbitrary object categories, which allows users to replace any target region with a reference object using only a coarse mask. To address the issues of severe homogenization and limited category coverage in existing datasets, we construct a large-scale multi-category dataset, UniEdit-500K, which includes 8 major categories, 209 fine-grained subcategories, and a total of 500,104 image pairs. Such rich category diversity poses new challenges for the model, requiring it to automatically learn semantic relationships and distinctions across categories. To this end, we introduce the Mixture of Transformer module, which performs differentiated modeling of various object categories through dynamic expert selection, and further enhances cross-category semantic transfer and generalization through collaboration among experts. In addition, we propose a Mask Annealing Training Strategy (MATS) that progressively relaxes mask precision during training, reducing the model's reliance on accurate masks and improving robustness across diverse editing tasks. Extensive experiments on benchmarks such as VITON-HD and AnyInsertion demonstrate that A^2-Edit consistently outperforms existing approaches across all metrics, providing a new and efficient solution for arbitrary object editing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。