用扩散模型去除真实视频中的阴影,保留细节并保持时间连贯性。
WildShadowRemover: In-the-Wild Video Shadow Removal via Detail-Preserving Video Diffusion Models

- 基于预训练扩散模型,通过LoRA微调实现视频去影。
- 引入高频纹理恢复模块,显著提升细节保真度。
- 适合需要高质量去影的视频编辑与影视制作场景。
真实世界视频去影因光照复杂、阴影形态多样及训练数据有限而极具挑战。尽管对视觉与图形应用至关重要,但在开放场景中仍研究不足。本文提出WildShadowRemover,通过LoRA微调预训练视频扩散模型,实现鲁棒去影。为在保留生成先验的同时增强细节,我们向冻结的VAE解码器添加细节注入模块,并引入阴影掩码引导的频域分解调制模块,以选择性恢复高频纹理并抑制阴影伪影。同时,利用Depth Anything 3提供的单目深度先验,在复杂光照下提供几何感知指导。我们还构建了大规模成对视频去影数据集WildShadow,涵盖多种合成场景。大量实验表明,该方法在去影质量与时间一致性上优于现有方法,生成视觉质量高、泛化能力强的无影视频。
原文摘要 · Abstract (English)
Video shadow removal in the wild remains challenging due to complex illumination, diverse shadow appearances, and limited training data. Despite its importance to numerous vision and graphics applications, it remains largely unexplored in unconstrained real-world scenarios. To address this gap, we present WildShadowRemover, a framework that adapts a pretrained video diffusion model for robust video shadow removal via LoRA fine-tuning. To preserve fine image details while retaining the model's powerful generative prior, we augment the frozen VAE decoder with a detail injection module and introduce a shadow-mask-guided frequency-decomposed modulation module to selectively restore high-frequency textures while suppressing shadow artifacts. Monocular depth priors from Depth Anything 3 further provide geometry-aware guidance under challenging lighting conditions. We also construct WildShadow, a large-scale paired video shadow removal dataset and benchmark, covering diverse synthetic scenes. Extensive experiments demonstrate that our method outperforms existing approaches in shadow removal quality and temporal consistency, producing temporally coherent shadow-free videos with superior visual quality and strong generalization across challenging in-the-wild scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。