无需相机位姿信息,实现物理真实且时序稳定的视频重光照。
Relit-LiVE: Relight Video by Jointly Learning Environment Video

- 联合学习环境视频与重光照结果,直接生成对齐视角的环境图。
- 在真实视频上显著减少畸变、材质破损和时间伪影。
- 适合需要动态光影和自由视角的视频编辑场景。
近期进展表明,大规模视频扩散模型可通过先分解视频为内在场景表示,再在新光照下进行前向渲染,用作神经渲染器。然而,该范式严重依赖准确的内在分解,而真实视频中的分解往往不可靠,常导致外观失真、材质断裂及累积的时间伪影。本文提出 Relit-LiVE,一种无需已知相机位姿即可生成物理一致、时序稳定的视频重光照的新框架。核心思路是将原始参考图像显式引入渲染过程,以恢复内在表示中丢失或损坏的关键场景线索。同时,提出一种新型环境视频预测形式,通过单一扩散过程同步生成重光照视频与逐帧对齐相机视角的环境图。该联合预测强化了几何-光照对齐,自然支持动态光照与相机运动,显著提升视频重光照的物理一致性,并降低对已知每帧相机位姿的要求。大量实验表明,Relit-LiVE 在合成与真实世界基准上持续优于现有最先进视频重光照与神经渲染方法。此外,本框架还可自然拓展至场景级渲染、材质编辑、物体插入与流媒体视频重光照等下游应用。项目地址:https://github.com/zhuxing0/Relit-LiVE。
原文摘要 · Abstract (English)
Recent advances have shown that large-scale video diffusion models can be repurposed as neural renderers by first decomposing videos into intrinsic scene representations and then performing forward rendering under novel illumination. While promising, this paradigm fundamentally relies on accurate intrinsic decomposition, which remains highly unreliable for real-world videos and often leads to distorted appearances, broken materials, and accumulated temporal artifacts during relighting. In this work, we present Relit-LiVE, a novel video relighting framework that produces physically consistent, temporally stable results without requiring prior knowledge of camera pose. Our key insight is to explicitly introduce raw reference images into the rendering process, enabling the model to recover critical scene cues that are inevitably lost or corrupted in intrinsic representations. Furthermore, we propose a novel environment video prediction formulation that simultaneously generates relit videos and per-frame environment maps aligned with each camera viewpoint in a single diffusion process. This joint prediction enforces strong geometric-illumination alignment and naturally supports dynamic lighting and camera motion, significantly improving physical consistency in video relighting while easing the requirement of known per-frame camera pose. Extensive experiments demonstrate that Relit-LiVE consistently outperforms state-of-the-art video relighting and neural rendering methods across synthetic and real-world benchmarks. Beyond relighting, our framework naturally supports a wide range of downstream applications, including scene-level rendering, material editing, object insertion, and streaming video relighting. The Project is available at https://github.com/zhuxing0/Relit-LiVE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。