arXiv:2602.01391cs.CV2026-02

用重光照测试视觉模型对物理世界的理解能力,发现语义抽象会损失光照细节。

Relighting as a Probe of Visual Priors via Augmented Latent Intrinsics

  • 通过生成式重光照框架,融合像素级特征与隐式光照模型
  • 在金属、透明材质上重光照质量提升显著,尤其改善光泽感还原
  • 适合研究视觉表征物理真实性或做生成模型评估的学者

图像到图像的重光照需要分离光照与场景属性,同时保留密集的几何、材质和光度线索。本文将该任务作为对视觉先验的探测:不同于强调不变性的识别任务,重光照考验特征是否保留光传输所需信息。通过一个受控的生成式重光照框架,我们发现强语义编码器反而会降低重光照质量,揭示出语义抽象与物理保真之间的权衡。为此提出增强型隐式内在(ALI)模型,通过融合密集像素对齐特征并利用无标签真实图像对进行自监督优化,实现性能平衡。ALI在高光、金属和透明材质上显著提升重光照效果,表明生成式重光照是量化视觉编码器对物理世界建模能力的有效工具。

原文摘要 · Abstract (English)

Image-to-image relighting requires representations that separate illumination from scene properties while preserving dense geometry, material, and photometric cues. We use this task as a probe of visual priors: unlike recognition tasks that reward invariance, relighting tests whether visual features retain the information needed for light transfer. Through a controlled generative relighting framework, we find that strong semantic encoders can degrade relighting quality, exposing a semantic--photometric trade-off between abstraction and physical fidelity. We introduce Augmented Latent Intrinsics (ALI), which balances this trade-off by fusing dense, pixel-aligned visual features into a latent-intrinsic relighting model and refining it with self-supervision on unlabeled real image pairs. ALI improves relighting quality, especially on glossy, metallic, and transparent materials, and demonstrates that generative relighting is an effective tool for quantifying what visual encoders encode about the physical world.

重光照视觉先验生成模型隐式表示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。