arXiv:2504.03011cs.CV2025-04CVPR被引 14

一次搞定任意人体的光影重制与背景融合,效果自然且时间连贯。

Comprehensive Relighting: Generalizable and Consistent Monocular Human Relighting and Harmonization

  • 用扩散模型做通用图像先验,分阶段实现光影重制与背景融合
  • 无监督学习真实视频中的光照周期一致性,提升时间连贯性
  • 适合需要高质量人像光影处理的影视与虚拟拍摄场景

本文提出 Comprehensive Relighting,首个能统一控制并融合任意场景中人体任意部位光影的全链路方法。由于缺乏数据集,现有基于图像的光影重制模型多局限于特定场景(如人脸或静止人体)。为此,我们复用预训练扩散模型作为通用图像先验,并在粗到精框架中联合建模人体光影重制与背景融合。为增强时序连贯性,引入无监督时序光照模型,从大量真实视频中学习光照周期一致性,无需真值标注。推理时,通过时空特征融合算法将时序光照模块与扩散模型结合,无需额外训练;并引入新型引导精修后处理,保留输入图像的高频细节。实验表明,Comprehensive Relighting 在泛化性和时序连贯性上均优于现有图像级人体光影重制与融合方法。

原文摘要 · Abstract (English)

This paper introduces Comprehensive Relighting, the first all-in-one approach that can both control and harmonize the lighting from an image or video of humans with arbitrary body parts from any scene. Building such a generalizable model is extremely challenging due to the lack of dataset, restricting existing image-based relighting models to a specific scenario (e.g., face or static human). To address this challenge, we repurpose a pre-trained diffusion model as a general image prior and jointly model the human relighting and background harmonization in the coarse-to-fine framework. To further enhance the temporal coherence of the relighting, we introduce an unsupervised temporal lighting model that learns the lighting cycle consistency from many real-world videos without any ground truth. In inference time, our temporal lighting module is combined with the diffusion models through the spatio-temporal feature blending algorithms without extra training; and we apply a new guided refinement as a post-processing to preserve the high-frequency details from the input image. In the experiments, Comprehensive Relighting shows a strong generalizability and lighting temporal coherence, outperforming existing image-based human relighting and harmonization methods.

光影重制扩散模型时序一致人体生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。