arXiv:2503.14485cs.GRcs.CV2025-03CVPR被引 31

用混合数据训练扩散模型,实现逼真且稳定的人脸光照重渲染。

Lux Post Facto: Learning Portrait Performance Relighting with Conditional Video Diffusion and a Hybrid Dataset

  • 基于预训练视频扩散模型,设计光照注入机制实现精准控制。
  • 在静态与真实场景视频混合数据上训练,无需配对光照视频。
  • 生成效果兼具真实感与时间一致性,适合影视与虚拟人应用。

视频人脸光照重渲染因需兼顾逼真度与时序稳定性而极具挑战性,通常依赖强大模型设计及高质量成对视频数据(如逐光拍摄的OLAT数据)。本文提出Lux Post Facto,一种新型人脸视频光照重渲染方法,可生成高保真且时序一致的光照效果。模型基于先进预训练视频扩散模型,引入新的光照注入机制以实现精确控制,利用空间与时间生成能力解决病态的重渲染问题。采用由静态表情OLAT数据与真实场景人物表演视频组成的混合数据集联合学习光照重渲染与时序建模,避免了获取不同光照条件下成对视频的需求。大量实验表明,该方法在逼真度与时序一致性方面均达到当前最优水平。

原文摘要 · Abstract (English)

Video portrait relighting remains challenging because the results need to be both photorealistic and temporally stable. This typically requires a strong model design that can capture complex facial reflections as well as intensive training on a high-quality paired video dataset, such as dynamic one-light-at-a-time (OLAT). In this work, we introduce Lux Post Facto, a novel portrait video relighting method that produces both photorealistic and temporally consistent lighting effects. From the model side, we design a new conditional video diffusion model built upon state-of-the-art pre-trained video diffusion model, alongside a new lighting injection mechanism to enable precise control. This way we leverage strong spatial and temporal generative capability to generate plausible solutions to the ill-posed relighting problem. Our technique uses a hybrid dataset consisting of static expression OLAT data and in-the-wild portrait performance videos to jointly learn relighting and temporal modeling. This avoids the need to acquire paired video data in different lighting conditions. Our extensive experiments show that our model produces state-of-the-art results both in terms of photorealism and temporal consistency.

视频生成扩散模型光照重渲染

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。