arXiv:2509.03680cs.GRcs.AI2025-09NeurIPS被引 17

用视频扩散模型从图像推算真实光照,精度超越现有方法

LuxDiT: Lighting Estimation with Video Diffusion Transformer

  • 用视频扩散变压器微调生成HDR环境图
  • 在真实场景上实现高保真光照重建,细节更逼真
  • 适合需要精准光照估计的3D渲染与虚拟现实应用

从单张图像或视频中估计场景光照是计算机视觉与图形学中的长期挑战。基于学习的方法受限于高质量HDR环境图的稀缺性,这类数据获取成本高且多样性不足。尽管近期生成模型提供了强大的图像合成先验,但光照估计仍因依赖间接视觉线索、需推断全局上下文以及恢复高动态范围输出而困难重重。我们提出LuxDiT,一种新的数据驱动方法,通过微调视频扩散变换器,根据视觉输入生成HDR环境图。模型在大规模合成数据集上训练,涵盖多种光照条件,能够从间接线索中推断照明信息,并有效泛化至真实场景。为提升输入与预测环境图之间的语义对齐,我们引入基于收集的HDR全景图数据集的低秩适应微调策略。该方法在定量与定性评估中均优于现有最先进技术,生成具有真实角度高频细节的准确光照预测。

原文摘要 · Abstract (English)

Estimating scene lighting from a single image or video remains a longstanding challenge in computer vision and graphics. Learning-based approaches are constrained by the scarcity of ground-truth HDR environment maps, which are expensive to capture and limited in diversity. While recent generative models offer strong priors for image synthesis, lighting estimation remains difficult due to its reliance on indirect visual cues, the need to infer global (non-local) context, and the recovery of high-dynamic-range outputs. We propose LuxDiT, a novel data-driven approach that fine-tunes a video diffusion transformer to generate HDR environment maps conditioned on visual input. Trained on a large synthetic dataset with diverse lighting conditions, our model learns to infer illumination from indirect visual cues and generalizes effectively to real-world scenes. To improve semantic alignment between the input and the predicted environment map, we introduce a low-rank adaptation finetuning strategy using a collected dataset of HDR panoramas. Our method produces accurate lighting predictions with realistic angular high-frequency details, outperforming existing state-of-the-art techniques in both quantitative and qualitative evaluations.

光照估计扩散模型视频生成HDR

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。