用残差生成控制扩散模型,精准还原阴影区域内容。
Controlling the Latent Diffusion Model for Generative Image Shadow Removal via Residual Generation
- 通过生成图像残差而非从零重建,保留原图细节。
- 在多个数据集上实现高保真阴影去除,优于现有方法。
- 适合需要精确内容还原的图像修复场景。
大规模生成模型在多种视觉任务中取得显著进展,但在图像去阴影方面仍面临挑战。现有方法常生成多样但缺乏保真度的细节,难以满足去阴影对内容精确还原的需求。本文提出利用扩散模型生成并优化图像残差,充分挖掘阴影图像中的内在细节信息,实现更高效、更忠实的无影内容重建。为防止生成过程中的误差累积,设计了跨时间步自增强训练策略,利用网络自身扩充训练数据,并动态修正生成轨迹,提升输出准确性与鲁棒性。此外,针对大模型编码解码过程中原始细节丢失的问题,构建了内容保持型编码器-解码器结构,结合控制机制与多尺度跳跃连接,实现高质量无影图像重建。实验表明,该方法基于大尺度潜在扩散先验,能有效生成高质量结果并忠实保留阴影区域原内容。
原文摘要 · Abstract (English)
Large-scale generative models have achieved remarkable advancements in various visual tasks, yet their application to shadow removal in images remains challenging. These models often generate diverse, realistic details without adequate focus on fidelity, failing to meet the crucial requirements of shadow removal, which necessitates precise preservation of image content. In contrast to prior approaches that aimed to regenerate shadow-free images from scratch, this paper utilizes diffusion models to generate and refine image residuals. This strategy fully uses the inherent detailed information within shadowed images, resulting in a more efficient and faithful reconstruction of shadow-free content. Additionally, to revent the accumulation of errors during the generation process, a crosstimestep self-enhancement training strategy is proposed. This strategy leverages the network itself to augment the training data, not only increasing the volume of data but also enabling the network to dynamically correct its generation trajectory, ensuring a more accurate and robust output. In addition, to address the loss of original details in the process of image encoding and decoding of large generative models, a content-preserved encoder-decoder structure is designed with a control mechanism and multi-scale skip connections to achieve high-fidelity shadow-free image reconstruction. Experimental results demonstrate that the proposed method can reproduce high-quality results based on a large latent diffusion prior and faithfully preserve the original contents in shadow regions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。