单视频一键重打光,效果逼真且稳定。
UniRelight: Learning Joint Decomposition and Synthesis for Video Relighting

- 联合估计材质与生成新光照,一步完成重打光。
- 在真实视频上训练,跨场景泛化能力强。
- 适合需要高质量光影还原的影视与动画制作。
我们解决单图或视频重打光的挑战,该任务需精确理解场景本质属性并合成高质量光传输效果。现有端到端模型受限于配对多光照数据稀缺,难以跨场景泛化;而两阶段逆向与正向渲染方法虽降低数据依赖,却易累积误差,复杂光照或复杂材质下生成结果不真实。本文提出通用性方法,通过视频扩散模型在单次前向传播中联合估计反照率并合成重打光输出,增强隐式场景理解,实现阴影、反射、透明等复杂材质交互的逼真还原。模型在合成多光照数据及大量自动标注的真实视频上训练,展现出强跨域泛化能力,在视觉保真度和时序一致性上均优于先前方法。
原文摘要 · Abstract (English)
We address the challenge of relighting a single image or video, a task that demands precise scene intrinsic understanding and high-quality light transport synthesis. Existing end-to-end relighting models are often limited by the scarcity of paired multi-illumination data, restricting their ability to generalize across diverse scenes. Conversely, two-stage pipelines that combine inverse and forward rendering can mitigate data requirements but are susceptible to error accumulation and often fail to produce realistic outputs under complex lighting conditions or with sophisticated materials. In this work, we introduce a general-purpose approach that jointly estimates albedo and synthesizes relit outputs in a single pass, harnessing the generative capabilities of video diffusion models. This joint formulation enhances implicit scene comprehension and facilitates the creation of realistic lighting effects and intricate material interactions, such as shadows, reflections, and transparency. Trained on synthetic multi-illumination data and extensive automatically labeled real-world videos, our model demonstrates strong generalization across diverse domains and surpasses previous methods in both visual fidelity and temporal consistency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。