无需训练即可精准插入物体,让新对象自然融入复杂场景。
Add-it: Training-Free Object Insertion in Images With Pretrained Diffusion Models
- 利用图像、文本与生成图三者信息扩展注意力机制
- 在真实与生成图像上均达当前最优,人类偏好超80%
- 适合希望快速实现高质量图像编辑的设计师与开发者
基于文本指令向图像中添加物体是语义图像编辑中的挑战性任务,需在保留原场景结构与自然融合新物体之间取得平衡。现有模型常难以在复杂场景中找到合适的放置位置。本文提出 Add-it,一种无需训练的方法,通过扩展扩散模型的注意力机制,融合场景图像、文本提示与生成图像三方面信息。其加权扩展注意力机制在保持结构一致性和细节的同时,确保物体放置自然合理。无需特定任务微调,Add-it 在真实图像与生成图像插入基准测试中均达到领先水平,包括我们新构建的“可放置性评估基准”(Additing Affordance Benchmark),优于监督方法。人工评估显示其在超过80%的情况下更受青睐,且在多种自动指标上均有提升。
原文摘要 · Abstract (English)
Adding Object into images based on text instructions is a challenging task in semantic image editing, requiring a balance between preserving the original scene and seamlessly integrating the new object in a fitting location. Despite extensive efforts, existing models often struggle with this balance, particularly with finding a natural location for adding an object in complex scenes. We introduce Add-it, a training-free approach that extends diffusion models' attention mechanisms to incorporate information from three key sources: the scene image, the text prompt, and the generated image itself. Our weighted extended-attention mechanism maintains structural consistency and fine details while ensuring natural object placement. Without task-specific fine-tuning, Add-it achieves state-of-the-art results on both real and generated image insertion benchmarks, including our newly constructed "Additing Affordance Benchmark" for evaluating object placement plausibility, outperforming supervised methods. Human evaluations show that Add-it is preferred in over 80% of cases, and it also demonstrates improvements in various automated metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。