无需微调,一键实现物体替换与场景和谐统一
Towards Source-Aware Object Swapping with Initial Noise Perturbation
- 通过初始噪声空间的频率分离扰动生成伪配对图像
- 实现零样本推理,保持物体与场景高保真度
- 适合需要快速物体替换的视觉编辑场景
物体替换旨在将场景中的源物体替换为参考物体,同时保持物体真实度、场景真实度和物体-场景和谐性。现有方法或需逐物体微调且推理缓慢,或依赖额外成对数据,这些数据大多仅在不同上下文中展示同一物体,迫使模型依赖背景线索而非学习跨物体对齐。我们提出 SourceSwap,一种自监督且源感知的框架,用于学习跨物体对齐。核心思路是通过初始噪声空间的频率分离扰动,从任意图像中合成高质量伪配对图像,改变外观但保留姿态、粗略形状和场景布局,无需视频、多视角数据或额外图像。我们训练一个带全源条件的双U-Net和无噪声参考编码器,实现直接跨物体对齐、无需微调的零样本推理以及轻量级迭代优化。此外,我们构建了 SourceBench,一个高分辨率、类别更丰富、交互更复杂的基准数据集。实验表明,SourceSwap 在保真度、场景保留和自然和谐方面表现更优,并可有效迁移至主体驱动优化和人脸替换等编辑任务。
原文摘要 · Abstract (English)
Object swapping aims to replace a source object in a scene with a reference object while preserving object fidelity, scene fidelity, and object-scene harmony. Existing methods either require per-object finetuning and slow inference or rely on extra paired data that mostly depict the same object across contexts, forcing models to rely on background cues rather than learning cross-object alignment. We propose SourceSwap, a self-supervised and source-aware framework that learns cross-object alignment. Our key insight is to synthesize high-quality pseudo pairs from any image via a frequency-separated perturbation in the initial-noise space, which alters appearance while preserving pose, coarse shape, and scene layout, requiring no videos, multi-view data, or additional images. We then train a dual U-Net with full-source conditioning and a noise-free reference encoder, enabling direct inter-object alignment, zero-shot inference without per-object finetuning, and lightweight iterative refinement. We further introduce SourceBench, a high-quality benchmark with higher resolution, more categories, and richer interactions. Experiments demonstrate that SourceSwap achieves superior fidelity, stronger scene preservation, and more natural harmony, and it transfers well to edits such as subject-driven refinement and face swapping.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。