arXiv:2508.17302cs.CV2025-08被引 1

无需训练即可精准插入指定物体,保持结构与外观一致

PosBridge: Multi-View Positional Embedding Transplant for Identity-Aware Image Editing

  • 通过位置嵌入移植引导扩散模型复现参考物体结构
  • 采用角中心布局输入,实现目标区域内容一致性生成
  • 高效省训,适合需要快速编辑的视觉创作场景

局部主体驱动图像编辑旨在无缝将用户指定对象融入目标场景。随着生成模型规模扩大,训练成本日益高昂,亟需无需训练且可扩展的编辑框架。为此,我们提出PosBridge——一种高效灵活的自定义物体插入方法。核心是位置嵌入移植,引导扩散模型忠实复现参考物体的结构特征。同时引入角中心布局,将参考图像与背景图拼接后输入FLUX.1-Fill模型。在逐步去噪过程中,位置嵌入移植用于引导目标区域的噪声分布向参考物体靠近。该方式有效促使FLUX.1-Fill模型在指定位置合成身份一致的内容。大量实验表明,PosBridge在结构一致性、外观保真度和计算效率上均优于主流基线,展现出实际应用价值与广泛推广潜力。

原文摘要 · Abstract (English)

Localized subject-driven image editing aims to seamlessly integrate user-specified objects into target scenes. As generative models continue to scale, training becomes increasingly costly in terms of memory and computation, highlighting the need for training-free and scalable editing frameworks. To this end, we propose PosBridge--an efficient and flexible framework for inserting custom objects. A key component of our method is positional embedding transplant, which guides the diffusion model to faithfully replicate the structural characteristics of reference objects. Meanwhile, we introduce the Corner Centered Layout, which concatenates reference images and the background image as input to the FLUX.1-Fill model. During progressive denoising, positional embedding transplant is applied to guide the noise distribution in the target region toward that of the reference object. In this way, Corner Centered Layout effectively directs the FLUX.1-Fill model to synthesize identity-consistent content at the desired location. Extensive experiments demonstrate that PosBridge outperforms mainstream baselines in structural consistency, appearance fidelity, and computational efficiency, showcasing its practical value and potential for broad adoption.

图像编辑扩散模型无训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。