让物体自适应融合进背景,实现更自然的图像合成。
DreamFuse: Adaptive Image Fusion with Diffusion Transformer
- 用扩散变换器+位置调制实现前景与背景的交互融合
- 通过人类反馈优化,提升背景一致性与前景和谐度
- 支持文本驱动的属性编辑,适用于创意设计场景
图像融合旨在将前景对象与背景场景无缝结合,生成真实且协调的融合图像。现有方法多直接将物体插入背景,而自适应、交互式融合仍具挑战性,需使前景能根据背景上下文调整自身。为此,我们提出一种人机协同的迭代数据生成流程,利用有限初始数据与多样文本提示,生成涵盖放置、手持、穿戴及风格迁移等多种情境的融合数据集。基于此,我们引入DreamFuse,一种基于扩散变换器(DiT)的新方法,可生成同时包含前景与背景信息的一致且协调的融合图像。DreamFuse采用位置仿射机制,将前景的大小和位置注入背景,通过共享注意力实现有效交互。此外,我们使用基于人类反馈的局部直接偏好优化来精炼DreamFuse,增强背景一致性与前景和谐性。实验表明,该方法在多项指标上优于当前最先进模型,且支持对融合结果的文本驱动属性编辑。
原文摘要 · Abstract (English)
Image fusion seeks to seamlessly integrate foreground objects with background scenes, producing realistic and harmonious fused images. Unlike existing methods that directly insert objects into the background, adaptive and interactive fusion remains a challenging yet appealing task. It requires the foreground to adjust or interact with the background context, enabling more coherent integration. To address this, we propose an iterative human-in-the-loop data generation pipeline, which leverages limited initial data with diverse textual prompts to generate fusion datasets across various scenarios and interactions, including placement, holding, wearing, and style transfer. Building on this, we introduce DreamFuse, a novel approach based on the Diffusion Transformer (DiT) model, to generate consistent and harmonious fused images with both foreground and background information. DreamFuse employs a Positional Affine mechanism to inject the size and position of the foreground into the background, enabling effective foreground-background interaction through shared attention. Furthermore, we apply Localized Direct Preference Optimization guided by human feedback to refine DreamFuse, enhancing background consistency and foreground harmony. DreamFuse achieves harmonious fusion while generalizing to text-driven attribute editing of the fused results. Experimental results demonstrate that our method outperforms state-of-the-art approaches across multiple metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。