arXiv:2605.28811cs.CV2026-05

让视频人像光照自然匹配背景,解决动态光影抖动问题

HarmoVid: Relightful Video Portrait Harmonization

论文配图:HarmoVid: Relightful Video Portrait Harmonization
图 1 · 摘自论文原文
  • 用去闪烁模型稳定视频光照,减少帧间跳变
  • 结合真实与合成视频训练,实现高保真光影调整
  • 适合需要自然光影融合的视频创作人群

我们提出一种视频人像光照协调方法,使前景视频的光照与目标背景场景一致,调整阴影、色调和光照强度(称为“重光”协调)。与图像不同,获取相同动作在不同光照下录制的配对视频数据在实践中不可行且无法扩展。虽然可将现有图像级协调模型逐帧应用于视频,但结果常出现显著时间抖动。为此,我们引入一种新型光照去闪烁模型,有效抑制全局与局部光照闪烁伪影。基于这些优化后的无闪烁数据,我们的视频扩散模型结合真实与合成视频进行训练,生成高质量视频协调结果。进一步提出非对称透明掩码条件机制,从真实视频中学习清晰边界。实验表明,本模型在时间连贯性、自然度、边界清晰度及物理合理的光照行为方面表现优异,同时相比以往图像与视频方法保持更强的重光表达能力。

原文摘要 · Abstract (English)

We present a method for harmonizing the lighting of a foreground video to match a target background scene, adjusting shadows, color tone, and illumination intensity (relightful harmonization). Unlike images, acquiring labeled data for videos, where identical motions are recorded under different lighting conditions, is practically infeasible and non-scalable. While one way to create such paired data is to apply existing image-based harmonization models frame by frame to a video, the resulting outputs often suffer from significant temporal jitters. We overcome this problem by introducing a novel lighting deflickering model that can stabilize the global and local lighting flickering artifacts. Our video diffusion model learns from these upgraded deflickered data with a volume of real and synthetic videos to generate high-quality video harmonization results. We further propose an asymmetric alpha mask conditioning technique to learn the clean boundaries from real videos. Experiments demonstrate that our model achieves strong temporal coherence, naturalness, cleaner boundaries, and physically meaningful lighting behavior, while maintaining strong relighting expressiveness compared to prior image-based and video-based harmonization methods.

视频生成光照协调扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。