用大模型实现视频光影一致重布景,保留主体原貌。
Lumen: Consistent Video Relighting and Harmonious Background Replacement with Video Generative Models
- 基于大模型与混合数据集,端到端实现光影同步重布景。
- 在真实与合成视频上训练,保持前景材质与光照一致性。
- 适合影视后期、虚拟拍摄等需要精准光影控制的场景。
视频重布光是一项挑战性高且实用的任务,旨在更换视频背景的同时,对前景光照进行协调调整并实现自然融合。过程中需保持前景原有属性(如反照率),并在时间帧间保持光照一致性。本文提出Lumen,一种基于大规模视频生成模型的端到端视频重布光框架,支持灵活文本指令控制光照与背景。针对高质量带标注的多光照条件视频稀缺问题,构建了一个包含真实与合成视频的大规模数据集:合成域利用社区丰富的3D资源,通过先进3D渲染引擎生成多样环境下的视频对;真实域则采用基于HDR的光照模拟,弥补野外配对视频缺失。依托该数据集,设计联合训练策略,有效发挥合成视频的物理一致性与真实视频的泛化分布优势。为此,在模型中引入领域感知适配器,解耦光照重布景与领域外观分布的学习。构建全面评估基准,从前景保持与视频一致性角度对比现有方法。实验表明,Lumen能有效将输入视频编辑为具有连贯光照与严格前景保持的电影级重布光视频。
原文摘要 · Abstract (English)
Video relighting is a challenging yet valuable task, aiming to replace the background in videos while correspondingly adjusting the lighting in the foreground with harmonious blending. During translation, it is essential to preserve the original properties of the foreground, e.g., albedo, and propagate consistent relighting among temporal frames. In this paper, we propose Lumen, an end-to-end video relighting framework developed on large-scale video generative models, receiving flexible textual description for instructing the control of lighting and background. Considering the scarcity of high-qualified paired videos with the same foreground in various lighting conditions, we construct a large-scale dataset with a mixture of realistic and synthetic videos. For the synthetic domain, benefiting from the abundant 3D assets in the community, we leverage advanced 3D rendering engine to curate video pairs in diverse environments. For the realistic domain, we adapt a HDR-based lighting simulation to complement the lack of paired in-the-wild videos. Powered by the aforementioned dataset, we design a joint training curriculum to effectively unleash the strengths of each domain, i.e., the physical consistency in synthetic videos, and the generalized domain distribution in realistic videos. To implement this, we inject a domain-aware adapter into the model to decouple the learning of relighting and domain appearance distribution. We construct a comprehensive benchmark to evaluate Lumen together with existing methods, from the perspectives of foreground preservation and video consistency assessment. Experimental results demonstrate that Lumen effectively edit the input into cinematic relighted videos with consistent lighting and strict foreground preservation. Our project page: https://lumen-relight.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。