用区域自适应注入提升图像融合细节与真实感。
Region-to-Region: Enhancing Generative Image Harmonization with Adaptive Regional Injection
- 通过区域间信息注入,保留前景高频细节
- 新数据集RPHarmony增强光照多样性,提升泛化能力
- 适合需要高精度图像融合的视觉生成任务
图像融合的目标是调整合成图像中的前景,使其与背景在视觉上保持一致。近年来,潜在扩散模型(LDM)被用于融合任务并取得显著成果,但依然存在细节丢失和融合能力有限的问题。现有合成数据集依赖色彩迁移,缺乏局部变化,无法捕捉复杂的现实光照条件。为此,本文提出区域到区域(Region-to-Region, R2R)变换方法,通过将合适区域的信息注入前景,既保留原始细节,又实现图像融合或生成新合成数据。为此,设计了Clear-VAE,利用自适应滤波保留前景高频细节并消除不协调成分;引入掩码感知的自适应通道注意力(MACA)的和谐控制器,动态根据前景与背景区域通道重要性调整融合过程。为解决数据集局限性,提出随机泊松混合(Random Poisson Blending),将合适区域的颜色与光照信息注入前景,构建更具多样性和挑战性的合成数据集RPHarmony。实验表明,该方法在定量指标和视觉一致性上均优于现有方法。此外,基于RPHarmony训练的模型在真实场景中生成更逼真的图像。代码、数据集及模型权重均已开源。
原文摘要 · Abstract (English)
The goal of image harmonization is to adjust the foreground in a composite image to achieve visual consistency with the background. Recently, latent diffusion model (LDM) are applied for harmonization, achieving remarkable results. However, LDM-based harmonization faces challenges in detail preservation and limited harmonization ability. Additionally, current synthetic datasets rely on color transfer, which lacks local variations and fails to capture complex real-world lighting conditions. To enhance harmonization capabilities, we propose the Region-to-Region transformation. By injecting information from appropriate regions into the foreground, this approach preserves original details while achieving image harmonization or, conversely, generating new composite data. From this perspective, We propose a novel model R2R. Specifically, we design Clear-VAE to preserve high-frequency details in the foreground using Adaptive Filter while eliminating disharmonious elements. To further enhance harmonization, we introduce the Harmony Controller with Mask-aware Adaptive Channel Attention (MACA), which dynamically adjusts the foreground based on the channel importance of both foreground and background regions. To address the limitation of existing datasets, we propose Random Poisson Blending, which transfers color and lighting information from a suitable region to the foreground, thereby generating more diverse and challenging synthetic images. Using this method, we construct a new synthetic dataset, RPHarmony. Experiments demonstrate the superiority of our method over other methods in both quantitative metrics and visual harmony. Moreover, our dataset helps the model generate more realistic images in real examples. Our code, dataset, and model weights have all been released for open access.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。