arXiv:2512.06358cs.CV2025-12被引 2

用新方法让生成模型看清图像中的反射层,提升去反射效果。

Rectifying Latent Space for Generative Single-Image Reflection Removal

  • 设计反射对称的变分自编码器,让模型理解反射的物理叠加原理。
  • 在多个基准上达到当前最佳性能,真实场景也表现良好。
  • 适合需要高精度图像修复的研究者与工程师。

单图去反射是一个高度病态的问题,现有方法难以准确解析被污染区域的构成,导致在真实场景中恢复失败且泛化能力差。本文将编辑用途的潜空间扩散模型重构为可有效感知和处理高度模糊、多层图像输入的结构,实现高质量输出。我们指出问题根源在于语义编码器的潜空间缺乏将复合图像视为各组成层线性叠加的内在结构。为此提出三个协同组件:反射等变的变分自编码器(VAE),使潜空间对齐反射形成的线性物理规律;可学习的任务特定文本嵌入,提供精准引导以绕过模糊语言描述;基于深度的早期分支采样策略,利用生成过程的随机性获得更优结果。大量实验表明,该模型在多个基准上达到新SOTA,并在挑战性真实场景中展现出良好泛化能力。

原文摘要 · Abstract (English)

Single-image reflection removal is a highly ill-posed problem, where existing methods struggle to reason about the composition of corrupted regions, causing them to fail at recovery and generalization in the wild. This work reframes an editing-purpose latent diffusion model to effectively perceive and process highly ambiguous, layered image inputs, yielding high-quality outputs. We argue that the challenge of this conversion stems from a critical yet overlooked issue, i.e., the latent space of semantic encoders lacks the inherent structure to interpret a composite image as a linear superposition of its constituent layers. Our approach is built on three synergistic components, including a reflection-equivariant VAE that aligns the latent space with the linear physics of reflection formation, a learnable task-specific text embedding for precise guidance that bypasses ambiguous language, and a depth-guided early-branching sampling strategy to harness generative stochasticity for promising results. Extensive experiments reveal that our model achieves new SOTA performance on multiple benchmarks and generalizes well to challenging real-world cases.

图像修复扩散模型反射去除

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。