arXiv:2603.15555cs.CV2026-03中稿 · CVPR

用稀疏物理线索引导扩散模型,实现单图光照的精准可控重布光。

Learning Latent Proxies for Controllable Single-Image Relighting

  • 通过少量PBR数据生成材料几何线索,构建轻量级潜空间代理编码器。
  • 在光照敏感区域施加掩码,使去噪过程聚焦于阴影相关像素,提升控制精度。
  • 适用于需要精细光照调节的场景,如数字内容创作与虚拟试穿。

单图重布光高度欠约束:微小光照变化可导致明暗、阴影和高光的剧烈非线性变化,而几何与材质信息未知。现有基于扩散的方法或依赖密集且脆弱的内在/深度缓冲区监督,或仅在潜空间操作缺乏物理依据,难以实现方向、强度、颜色的细粒度控制。本文发现完整内在分解对准确重布光并不必要。相反,少量但具有物理意义的提示——指示光照应如何改变及材质如何响应——已足够引导扩散模型。基于此,提出LightCtrl,其在两个层面融入物理先验:少样本潜空间代理编码器从有限的PBR监督中提取紧凑的材质-几何线索;光照感知掩码识别敏感光照区域,引导去噪器聚焦于与阴影相关的像素。为弥补稀缺的PBR数据,采用基于DPO的目标优化代理分支,确保预测线索的物理一致性。同时构建了大规模物体级数据集ScaLight,包含系统化变化的光照与完整的相机-光源元数据,支持物理一致且可控的训练。在物体与场景级基准上,本方法实现了高保真的光度重布光,并具备精确连续控制能力,超越先前扩散与内在基线方法,在受控光照偏移下最高提升2.4 dB PSNR,RMSE降低35%。

原文摘要 · Abstract (English)

Single-image relighting is highly under-constrained: small illumination changes can produce large, nonlinear variations in shading, shadows, and specularities, while geometry and materials remain unobserved. Existing diffusion-based approaches either rely on intrinsic or G-buffer pipelines that require dense and fragile supervision, or operate purely in latent space without physical grounding, making fine-grained control of direction, intensity, and color unreliable. We observe that a full intrinsic decomposition is unnecessary and redundant for accurate relighting. Instead, sparse but physically meaningful cues, indicating where illumination should change and how materials should respond, are sufficient to guide a diffusion model. Based on this insight, we introduce LightCtrl that integrates physical priors at two levels: a few-shot latent proxy encoder that extracts compact material-geometry cues from limited PBR supervision, and a lighting-aware mask that identifies sensitive illumination regions and steers the denoiser toward shading relevant pixels. To compensate for scarce PBR data, we refine the proxy branch using a DPO-based objective that enforces physical consistency in the predicted cues. We also present ScaLight, a large-scale object-level dataset with systematically varied illumination and complete camera-light metadata, enabling physically consistent and controllable training. Across object and scene level benchmarks, our method achieves photometrically faithful relighting with accurate continuous control, surpassing prior diffusion and intrinsic-based baselines, including gains of up to +2.4 dB PSNR and 35% lower RMSE under controlled lighting shifts.

重布光扩散模型可控生成物理先验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。