通过图像扭曲提升图像翻译中的细节保留能力
WarpI2I: Image Warping for Image-to-Image Translation

- 用显著性引导的扭曲-逆扭曲框架,将空间信息集中到关键区域
- 在人类光照和驾驶场景任务中,结构保真度和光影准确性显著提升
- 无需修改模型架构,计算开销极小,适合视频应用
图像到图像(I2I)翻译在人物光照重演和驾驶场景转换等任务中,已通过潜在扩散模型(LDMs)取得显著进展。然而,紧凑型LDM在处理高分辨率输入时,因编码器将数据压缩至低分辨率潜在空间,常导致细粒度结构丢失。为此,本文提出一种简单且基于显著性的扭曲-逆扭曲框架:在编码前将空间表示重新分配至显著区域,从而在不增加潜在分辨率的前提下增强结构细节保留。扭曲后的图像由原扩散模型处理,再通过逆扭曲映射回原始空间。此外,我们设计了一种基于外绘(outpainting)的合成数据生成流水线,可高效生成高质量成对光照重演数据。方法具备模型无关性,无需架构修改,计算开销极低。在人物光照、驾驶场景光照及图像转换任务上的实验表明,该方法在结构保真度、光照忠实性和图像质量方面均有提升,且可通过逐帧应用自然扩展至视频,保持良好时间稳定性。
原文摘要 · Abstract (English)
Image-to-image (I2I) translation has achieved strong results in tasks like human relighting and driving scene translation using latent diffusion models (LDMs). However, compact LDMs often struggle to preserve fine-grained structures because the encoder compresses high-resolution inputs into a spatially downsampled latent space. To address this issue, we propose a simple saliency-guided warp-unwarp framework that reallocates spatial representation toward salient regions before encoding, enabling better preservation of structural details without increasing latent resolution. The warped image is processed by the original diffusion model and then mapped back via an inverse warp. In addition, we propose a simple and efficient outpainting-based synthetic data generation pipeline to produce high-quality paired data for image relighting. Our method is model-agnostic, requires no architectural modification, and introduces negligible computational overhead. Experiments on human relighting, driving scene relighting, and translation demonstrate improved structural preservation, lighting faithfulness, and image quality, with our framework extending naturally to video via frame-by-frame application with good temporal stability. Project Webpage: https://shenzheng2000.github.io/WarpI2I.github.io
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。