用一张参考图,精准插入物体到任意场景中。
Insert Anything: Image Insertion via In-Context Editing in DiT
- 利用DiT模型的多模态注意力,支持掩码和文本双重控制。
- 在12万组数据上训练,可零样本适配多种插入任务。
- 适合创意设计、虚拟试衣等需要高保真融合的场景。
本文提出Insert Anything,一种基于参考图像的统一图像插入框架,可在用户指定的控制引导下,将参考图像中的物体无缝融入目标场景。该方法仅需在包含12万组提示-图像对的新数据集AnyInsertion上训练一次,即可泛化至人物、物体、服装等多种插入任务。该任务要求同时保留对象的身份特征与精细细节,并支持风格、颜色、纹理等灵活局部调整。为此,我们利用扩散变换器(DiT)的多模态注意力机制,实现掩码与文本双控编辑;并引入上下文编辑机制,将参考图像作为上下文信息,通过两种提示策略使插入元素与目标场景协调一致,同时忠实保留其独特特征。在AnyInsertion、DreamBooth和VTON-HD基准上的大量实验表明,本方法持续优于现有方法,展现出在创意内容生成、虚拟试穿和场景合成等实际应用中的巨大潜力。
原文摘要 · Abstract (English)
This work presents Insert Anything, a unified framework for reference-based image insertion that seamlessly integrates objects from reference images into target scenes under flexible, user-specified control guidance. Instead of training separate models for individual tasks, our approach is trained once on our new AnyInsertion dataset--comprising 120K prompt-image pairs covering diverse tasks such as person, object, and garment insertion--and effortlessly generalizes to a wide range of insertion scenarios. Such a challenging setting requires capturing both identity features and fine-grained details, while allowing versatile local adaptations in style, color, and texture. To this end, we propose to leverage the multimodal attention of the Diffusion Transformer (DiT) to support both mask- and text-guided editing. Furthermore, we introduce an in-context editing mechanism that treats the reference image as contextual information, employing two prompting strategies to harmonize the inserted elements with the target scene while faithfully preserving their distinctive features. Extensive experiments on AnyInsertion, DreamBooth, and VTON-HD benchmarks demonstrate that our method consistently outperforms existing alternatives, underscoring its great potential in real-world applications such as creative content generation, virtual try-on, and scene composition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。