arXiv:2607.04674cs.CV2026-07

用视频生成模型自动推断真实场景光照,无需额外训练。

Video Generation Models Are Inherent Lighting Estimators

论文配图:Video Generation Models Are Inherent Lighting Estimators
图 1 · 摘自论文原文
  • 把光照估计转为引导填图任务,插入反光球逼模型学物理反射。
  • 生成连续一致的HDR环境图,时间上稳定且符合真实光照变化。
  • 适合影视特效、3D重建方向,无需复杂标注即可用现成模型。

从单个自然场景视频中恢复动态环境光图对真实感渲染至关重要,但仍是难题。最近的视频生成模型能生成具有复杂光照的逼真场景,具备内在的光照理解能力。本文提出V-LITE(Video generation models are inherent lighting estimators)框架,通过将光照估计重构为引导视频修复任务,激活模型内部知识。受影视特效行业启发,我们在场景中插入一个合成的铬球,迫使模型根据时空上下文生成物理上合理的反射。为解决低动态范围(LDR)模型与高动态范围(HDR)域之间的差距,我们设计了适配HDR的变分自编码器(VAE),并采用高效的基于LoRA的微调策略。我们构建了一个混合数据集,包含高保真HDR图像提供真实HDR先验,以及真实世界中的HDR视频提供动态时空上下文。大量实验表明,V-LITE可生成时间连贯的HDR环境图,揭示现代视频扩散模型不仅是内容合成器,更是强大且内嵌的物理场景光照估计算器。

原文摘要 · Abstract (English)

Recovering dynamic environment maps from a single in-the-wild video is crucial for photorealistic rendering, yet remains a challenge. Recent video generation models can produce photorealistic scenes with complex lighting, possessing an inherent understanding of lighting. In this paper, we introduce V-LITE (Video generation models are inherent lighting estimators), a framework that unlocks this internal knowledge by reframing lighting estimation as a guided video inpainting task. Inspired by VFX industry practices, we insert a synthetic chrome ball into the scene to compel the model to generate physically plausible reflections from the surrounding spatio-temporal context. To bridge the gap from LDR-native models to the HDR domain, we design an HDR-aware VAE and employ an efficient LoRA-based fine-tuning strategy. We then construct a mixed dataset comprising high-fidelity HDR images to provide realistic HDR priors, and in-the-wild HDR videos to provide dynamic spatio-temporal context. Extensive experiments demonstrate that V-LITE produces temporally coherent HDR environment maps, revealing that modern video diffusion models are not merely synthesizers but also powerful, inherently capable estimators of physical scene lighting.

视频生成光照估计扩散模型HDR重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。