无需训练即可高保真迁移图像外观,保持场景结构完整。
A training-free framework for high-fidelity appearance transfer via diffusion transformers
- 分离结构与外观,通过注意力共享动态融合参考图像特征。
- 1024分辨率下实现高保真外观迁移,优于专用方法。
- 适合需要精确外观控制的图像编辑场景,如材质替换。
扩散变换器(DiTs)在生成任务中表现优异,但其全局自注意力机制使得基于参考图的可控编辑成为挑战。与U-Net不同,直接注入局部外观会破坏整体场景结构。为此,我们提出首个无需训练的框架,专为高保真外观迁移设计。核心是解耦结构与外观的协同系统:利用高保真反演建立源图像丰富的内容先验,捕捉光照与微纹理;创新的注意力共享机制则动态融合来自参考图的净化外观特征,受几何先验引导。该统一方法在1024像素分辨率下运行,涵盖语义属性迁移到细粒度材质应用等多种任务,实验表明其在结构保留与外观保真度上均达当前最优水平。
原文摘要 · Abstract (English)
Diffusion Transformers (DiTs) excel at generation, but their global self-attention makes controllable, reference-image-based editing a distinct challenge. Unlike U-Nets, naively injecting local appearance into a DiT can disrupt its holistic scene structure. We address this by proposing the first training-free framework specifically designed to tame DiTs for high-fidelity appearance transfer. Our core is a synergistic system that disentangles structure and appearance. We leverage high-fidelity inversion to establish a rich content prior for the source image, capturing its lighting and micro-textures. A novel attention-sharing mechanism then dynamically fuses purified appearance features from a reference, guided by geometric priors. Our unified approach operates at 1024px and outperforms specialized methods on tasks ranging from semantic attribute transfer to fine-grained material application. Extensive experiments confirm our state-of-the-art performance in both structural preservation and appearance fidelity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。