让单图灯光可调,能精准控制位置、颜色和亮度。
LightMover: Generative Light Movement with Color and Intensity Controls
- 用视频扩散先验建模光照变化,统一控制光的位置与外观。
- 实现41%控制序列压缩,保持编辑精度与光影真实感。
- 适合需要精细光照编辑的影视、游戏与设计场景。
我们提出LightMover,一个在单张图像上实现可控光照操作的框架,利用视频扩散先验生成物理合理的光照变化,无需重新渲染场景。将光照编辑建模为视觉标记空间中的序列到序列预测:给定一张图像和光照控制标记,模型同步调整光源位置、颜色与强度,以及由此产生的反射、阴影和衰减效果。这种对空间(移动)与外观(颜色、强度)的统一处理提升了操作精度与光照理解能力。我们进一步引入自适应标记剪枝机制,保留空间相关信息标记,紧凑编码非空间属性,使控制序列长度减少41%,同时保持编辑保真度。为训练该框架,我们构建了一个可扩展的渲染流水线,可在不同光照位置、颜色和强度下生成大量图像对,同时保持场景内容与原图一致。LightMover实现了对光照位置、颜色和强度的精确独立控制,在多种任务中达到高PSNR值,并展现出强大的语义一致性(通过DINO、CLIP评估)。
原文摘要 · Abstract (English)
We present LightMover, a framework for controllable light manipulation in single images that leverages video diffusion priors to produce physically plausible illumination changes without re-rendering the scene. We formulate light editing as a sequence-to-sequence prediction problem in visual token space: given an image and light-control tokens, the model adjusts light position, color, and intensity together with resulting reflections, shadows, and falloff from a single view. This unified treatment of spatial (movement) and appearance (color, intensity) controls improves both manipulation and illumination understanding. We further introduce an adaptive token-pruning mechanism that preserves spatially informative tokens while compactly encoding non-spatial attributes, reducing control sequence length by 41% while maintaining editing fidelity. To train our framework, we construct a scalable rendering pipeline that generates large numbers of image pairs across varied light positions, colors, and intensities while keeping the scene content consistent with the original image. LightMover enables precise, independent control over light position, color, and intensity, and achieves high PSNR and strong semantic consistency (DINO, CLIP) across different tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。