arXiv:2512.05000cs.CVcs.AI2025-12被引 1

用扩散Transformer和物理渲染数据,高效去除单张图像的反射

Reflection Removal through Efficient Adaptation of Diffusion Transformers

  • 复用预训练扩散Transformer,通过条件输入引导去反射
  • 合成真实玻璃材质数据,配合LoRA微调达到顶尖效果
  • 适合需要高保真去反射且无标注数据的场景

我们提出一种基于扩散变换器(DiT)的单图像反射去除框架,利用基础扩散模型在修复任务中的泛化能力。不依赖特定架构,而是通过将预训练的DiT模型条件于含反射输入,并引导其生成无反射传输层。系统分析了现有反射去除数据集的多样性、可扩展性和逼真度。为解决数据不足问题,构建基于Blender的物理基础渲染(PBR)管线,采用原理性双向散射分布函数(Principled BSDF)合成真实玻璃材质与反射效果。结合所提合成数据,使用高效的LoRA方式对基础模型进行适配,在域内和零样本基准测试中均取得领先性能。结果表明,当预训练扩散变压器与物理驱动的数据合成及高效适配结合时,可提供可扩展且高保真的反射去除方案。

原文摘要 · Abstract (English)

We introduce a diffusion-transformer (DiT) framework for single-image reflection removal that leverages the generalization strengths of foundation diffusion models in the restoration setting. Rather than relying on task-specific architectures, we repurpose a pre-trained DiT-based foundation model by conditioning it on reflection-contaminated inputs and guiding it toward clean transmission layers. We systematically analyze existing reflection removal data sources for diversity, scalability, and photorealism. To address the shortage of suitable data, we construct a physically based rendering (PBR) pipeline in Blender, built around the Principled BSDF, to synthesize realistic glass materials and reflection effects. Efficient LoRA-based adaptation of the foundation model, combined with the proposed synthetic data, achieves state-of-the-art performance on in-domain and zero-shot benchmarks. These results demonstrate that pretrained diffusion transformers, when paired with physically grounded data synthesis and efficient adaptation, offer a scalable and high-fidelity solution for reflection removal. Project page: https://hf.co/spaces/huawei-bayerlab/windowseat-reflection-removal-web

扩散模型去反射数据合成LoRA

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。