arXiv:2603.14241cs.CV2026-03

单图生成可控视角与光照的视频,统一建模更高效。

CamLit: Unified Video Diffusion with Explicit Camera and Lighting Control

  • 用一张图+相机轨迹+环境光,统一生成新视角和新光照视频。
  • 生成视频时序连贯、空间对齐,支持光照与视角同步控制。
  • 无需多模型切换,适合影视特效与虚拟拍摄场景使用。

我们提出 CamLit,首个统一的视频扩散模型,可从单张参考图像中联合实现新视角合成(NVS)与重光照。给定一张参考图像、用户定义的相机轨迹和环境贴图,CamLit 能生成该场景在新视角下指定光照条件的视频。在一个生成过程中,模型输出时间连贯且空间对齐的结果,包括重光照的新视角帧与对应的反照率帧,实现对相机姿态和光照的高质量控制。定性和定量实验表明,CamLit 在新视角合成与重光照任务上均达到当前最优水平,且不牺牲任一任务的视觉质量。结果证明,单一生成模型可有效整合相机与光照控制,简化视频生成流程,同时保持竞争力和一致的逼真度。

原文摘要 · Abstract (English)

We present CamLit, the first unified video diffusion model that jointly performs novel view synthesis (NVS) and relighting from a single input image. Given one reference image, a user-defined camera trajectory, and an environment map, CamLit synthesizes a video of the scene from new viewpoints under the specified illumination. Within a single generative process, our model produces temporally coherent and spatially aligned outputs, including relit novel-view frames and corresponding albedo frames, enabling high-quality control of both camera pose and lighting. Qualitative and quantitative experiments demonstrate that CamLit achieves high-fidelity outputs on par with state-of-the-art methods in both novel view synthesis and relighting, without sacrificing visual quality in either task. We show that a single generative model can effectively integrate camera and lighting control, simplifying the video generation pipeline while maintaining competitive performance and consistent realism.

视频生成扩散模型光照控制视角合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。