提出GENIE框架,实现参考图像的精准实例编辑。
Borrowing from anything: A generalizable framework for reference-guided instance editing
- 通过空间对齐与自适应残差缩放,分离参考图的内在外观与外在属性。
- 在AnyInsertion数据集上达到最新最高保真度与鲁棒性。
- 适合需要精确控制图像修改内容的研究者使用。
参考引导的实例编辑受限于语义纠缠问题,即参考图像的固有外观与其外部属性相互交织。核心挑战在于判断应从参考图像中借用哪些信息,并如何恰当地将其应用到目标图像上。为此,我们提出GENIE——一种可泛化的实例编辑框架,能够实现显式的解耦。GENIE首先通过空间对齐模块(SAM)纠正空间错位;接着,自适应残差缩放模块(ARSM)通过增强显著的内在特征并抑制外在属性来学习应借用的内容;最后,渐进式注意力融合(PAF)机制学习如何将该外观渲染到目标图像上,同时保持其结构完整性。在具有挑战性的AnyInsertion数据集上的大量实验表明,GENIE在保真度和鲁棒性方面均达到当前最优水平,为基于解耦的实例编辑树立了新标准。
原文摘要 · Abstract (English)
Reference-guided instance editing is fundamentally limited by semantic entanglement, where a reference's intrinsic appearance is intertwined with its extrinsic attributes. The key challenge lies in disentangling what information should be borrowed from the reference, and determining how to apply it appropriately to the target. To tackle this challenge, we propose GENIE, a Generalizable Instance Editing framework capable of achieving explicit disentanglement. GENIE first corrects spatial misalignments with a Spatial Alignment Module (SAM). Then, an Adaptive Residual Scaling Module (ARSM) learns what to borrow by amplifying salient intrinsic cues while suppressing extrinsic attributes, while a Progressive Attention Fusion (PAF) mechanism learns how to render this appearance onto the target, preserving its structure. Extensive experiments on the challenging AnyInsertion dataset demonstrate that GENIE achieves state-of-the-art fidelity and robustness, setting a new standard for disentanglement-based instance editing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。