用物理启发的扩散模型实现真实感全图重光照,提升泛化能力。
PI-Light: Physics-Inspired Diffusion for Full-Image Relighting
- 分两阶段设计,融合物理约束与注意力机制提升一致性
- 在多种材质上生成真实镜面高光和漫反射,优于现有方法
- 适用于真实图像编辑,尤其适合需要物理一致性的场景
全图重光照因缺乏大规模结构化配对数据、难以保证物理合理性,以及数据驱动先验导致的泛化能力有限而面临挑战。现有方法在合成到真实场景间的差距仍不理想。为此,我们提出物理启发的扩散模型PI-Light,采用两阶段框架:引入批次感知注意力以提升多图内在预测的一致性;设计物理引导的神经渲染模块,强制符合物理光传输规律;使用物理启发的损失函数,使训练过程趋于物理合理空间,增强对真实世界图像编辑的泛化能力;并构建一个在受控光照下采集的多样化物体与场景数据集。这些组件支持预训练扩散模型的高效微调,并提供下游评估基准。实验表明,PI-Light能有效生成多种材质的镜面高光与漫反射,相比以往方法展现出更强的真实场景泛化能力。
原文摘要 · Abstract (English)
Full-image relighting remains a challenging problem due to the difficulty of collecting large-scale structured paired data, the difficulty of maintaining physical plausibility, and the limited generalizability imposed by data-driven priors. Existing attempts to bridge the synthetic-to-real gap for full-scene relighting remain suboptimal. To tackle these challenges, we introduce Physics-Inspired diffusion for full-image reLight ($π$-Light, or PI-Light), a two-stage framework that leverages physics-inspired diffusion models. Our design incorporates (i) batch-aware attention, which improves the consistency of intrinsic predictions across a collection of images, (ii) a physics-guided neural rendering module that enforces physically plausible light transport, (iii) physics-inspired losses that regularize training dynamics toward a physically meaningful landscape, thereby enhancing generalizability to real-world image editing, and (iv) a carefully curated dataset of diverse objects and scenes captured under controlled lighting conditions. Together, these components enable efficient finetuning of pretrained diffusion models while also providing a solid benchmark for downstream evaluation. Experiments demonstrate that $π$-Light synthesizes specular highlights and diffuse reflections across a wide variety of materials, achieving superior generalization to real-world scenes compared with prior approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。