arXiv:2501.16330cs.CVcs.AI2025-01被引 31

RelightVid实现视频任意光照重置,保持时序一致性和材质真实性。

RelightVid: Temporal-Consistent Diffusion Model for Video Relighting

  • 基于扩散模型,支持背景视频、文本或环境图作为光照条件输入。
  • 在极端动态光照下训练,实现高时序一致性与真实光照先验保留。
  • 无需分解材质,适用于影视后期、虚拟拍摄等需要真实感光照的场景。

扩散模型在图像生成与编辑中表现卓越,近期进展已实现保持材质不变的图像光照重置。然而,视频光照重置因缺乏成对数据集、对输出保真度和时序一致性要求高,且扩散模型固有随机性而面临挑战。为此,我们提出RelightVid,一种灵活的视频光照重置框架,可接受背景视频、文本提示或环境图作为光照条件。该模型在真实场景视频上训练,结合精心设计的光照增强与极端动态光照下的渲染视频,实现无需内在分解的任意视频光照重置,同时保持高时序一致性并保留图像骨干的光照先验。

原文摘要 · Abstract (English)

Diffusion models have demonstrated remarkable success in image generation and editing, with recent advancements enabling albedo-preserving image relighting. However, applying these models to video relighting remains challenging due to the lack of paired video relighting datasets and the high demands for output fidelity and temporal consistency, further complicated by the inherent randomness of diffusion models. To address these challenges, we introduce RelightVid, a flexible framework for video relighting that can accept background video, text prompts, or environment maps as relighting conditions. Trained on in-the-wild videos with carefully designed illumination augmentations and rendered videos under extreme dynamic lighting, RelightVid achieves arbitrary video relighting with high temporal consistency without intrinsic decomposition while preserving the illumination priors of its image backbone.

视频重光照扩散模型时序一致性图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。