用扩散模型实现动态人像视频真实光照重演,保持表情与细节一致。
Pixel Cube: Diffusion-based Portrait Video Relighting Through Realistic Lighting Reproduction

- 基于真实采集与渲染数据训练,用HDR环境图控制光照。
- 生成视频光照逼真、时序连贯,保留皮肤纹路与胡须等细节。
- 支持未见人物、动作与光照条件,适合影视与摄影应用。
我们提出一种基于扩散模型的动态人像视频光照重演方法,实现照片级真实感与时间一致性。通过自建LED光照系统,采集真实动态人像视频与对应已知光照条件的数据,构建混合训练集。利用预训练视频扩散模型中的图像先验,并以每帧高动态范围(HDR)环境图作为光照控制信号,训练出高性能生成模型。此外,通过合成背景图控制相机曝光与色彩基调。模型可生成在新环境下真实和谐、时序稳定的重光人像视频,忠实保留表情、肤色、皱纹与面部毛发等精细特征。在多种未见主体、动作与光照条件下均表现良好。大量实验验证了其在野外视频重光任务中的优越性能,且在人像摄影中有实际应用价值。
原文摘要 · Abstract (English)
We present a diffusion-based method for relighting dynamic portrait videos with photorealism and temporal consistency. Our method is fueled by a hybrid training dataset that consists of real-captured and rendered dynamic portrait videos with diverse subject appearances, facial motions, head poses, and known lighting conditions. Specifically, we construct an LED-based lighting system for realistic lighting emulation and high-speed video relighting data acquisition. By leveraging the image priors embedded in pre-trained video diffusion models, and using per-frame high dynamic range (HDR) environment map as lighting control, we train a high-performance generative model for realistic and identity-preserving dynamic portrait video relighting. In addition to the environment map control, our model uses a synthesized background image to enable control on the camera's exposure level and color tone. Our model can produce temporally consistent relit portrait video that looks realistic and harmonious under a provided new environment and faithfully preserve the subject's expression and fine facial features, including skin tone, wrinkles, and facial hair. Our model generalizes well to unseen data, in terms of the subject appearance, motion, and lighting condition. We perform extensive experiments on relighting in-the-wild videos with various environment maps and demonstrate practical applications on portrait photography. Results show that our method achieves state-of-the-art performance in photorealism, lighting harmony, and temporal consistency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。