arXiv:2608.11562cs.CVcs.AI2026-08

首个基于扩散模型的视频去反射方法,实现单步去反射并支持真实评估。

From Synthesis to Removal: Physics-Grounded Reflection Simulation and Diffusion-Based Video Dereflection

论文配图:From Synthesis to Removal: Physics-Grounded Reflection Simulation and Diffusion-Based Video Dereflection
图 1 · 摘自论文原文
  • 通过物理建模生成带反射与无反射视频对,模拟玻璃粗糙、厚度等真实效应。
  • 提出S2R-Removal模型,单步去噪恢复清晰画面,速度优于非扩散基线。
  • 构建首个视频去反射基准S2R-Bench,支持客观与主观评价,推动领域发展。

通过玻璃拍摄的视频常含反射,影响视觉质量并干扰下游任务。尽管单图去反射研究充分,视频去反射因缺乏成对数据、时序一致性模型及专用评估基准而研究不足。本文提出闭环框架:S2R-Synthesis通过结构空间物理增强与训练过的视频扩散渲染器生成真实反射视频,模拟粗糙度导致的模糊、厚度引起的鬼影及反射率变化;基于合成数据,提出首个扩散模型驱动的视频去反射方法S2R-Removal,通过反射感知隐空间适配与单步像素几何精修,在一步去噪中恢复干净透射图像;进一步构建首个视频去反射评估基准S2R-Bench,支持全参考评估与真实人类感知测试。在S2R-Bench及多个公开图像基准上的实验表明,该方法性能达最先进水平,推理速度也快于部分非扩散基线,并验证了S2R-Synthesis的有效性。

原文摘要 · Abstract (English)

Videos captured through glass often contain reflections that degrade visual quality and interfere with downstream vision tasks. Although single-image reflection removal has been extensively studied, video reflection removal remains largely underexplored due to the lack of paired video data, temporally coherent removal models, and dedicated evaluation benchmarks. We present a closed-loop framework that unifies physics-grounded reflection simulation, diffusion-based video dereflection, and benchmark evaluation. Our S2R-Synthesis pipeline generates paired reflected and reflection-free videos by performing physics-grounded augmentation in the structure space and rendering realistic reflected videos with a trained video diffusion renderer; the augmentation models key glass-related effects including roughness-induced blur, thickness-induced ghosting, and reflectance variation. Based on the synthesized data, we introduce S2R-Removal, the first diffusion-based video reflection removal model, which adapts a pretrained video diffusion prior through reflection-aware latent adaptation and one-step pixel-geometric refinement, recovering the clean transmission in a single denoising step. We further build S2R-Bench, the first benchmark for video reflection removal, supporting both full-reference evaluation and real-world human perceptual assessment. Experiments on S2R-Bench and multiple public image benchmarks demonstrate state-of-the-art performance and faster inference than even non-diffusion baselines, and validate the effectiveness of S2R-Synthesis. Project page: https://codingwzp.github.io/VideoDereflection_S2R.

视频去反射扩散模型物理仿真评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。