arXiv:2410.05525cs.CV2024-10被引 12

用扩散模型生成无阴影人像,保留细节和真实光照。

Generative Portrait Shadow Removal

  • 将去阴影视为生成任务,全局重建人脸外观。
  • 在真实与合成数据上训练,实现自然去阴影效果。
  • 适合需要高质量人像修复的视觉应用。

我们提出一种高保真人像去阴影模型,通过预测受阴影和高光干扰下的人像外观来增强图像。人像去阴影是一个高度病态问题,单幅图像可能对应多种合理解。现有方法通过预测局部阴影传播残差解决此问题,但常不完整且对硬阴影产生不自然结果。我们通过将去阴影问题建模为生成任务,利用扩散模型以输入人像为条件,从头全局重建人体外观,克服了局部传播方法的局限性。为实现鲁棒自然的去阴影效果,提出组合式微调框架:先在背景融合数据集上微调预训练文本引导生成模型,使其协调前景与背景光照色彩;再在阴影配对数据集上进一步微调,生成无阴影人像。为弥补潜在扩散模型丢失的高频细节(如皱纹、斑点),引入引导上采样网络,从输入图中恢复原始细节。为支持该训练框架,我们使用光场捕捉系统和合成图形仿真构建了一个大规模高保真人像数据集。所提生成框架能有效去除自遮挡与外部遮挡引起的阴影,同时保持原有光照分布与高频细节,在真实环境采集的多样人物上也表现出强鲁棒性。

原文摘要 · Abstract (English)

We introduce a high-fidelity portrait shadow removal model that can effectively enhance the image of a portrait by predicting its appearance under disturbing shadows and highlights. Portrait shadow removal is a highly ill-posed problem where multiple plausible solutions can be found based on a single image. While existing works have solved this problem by predicting the appearance residuals that can propagate local shadow distribution, such methods are often incomplete and lead to unnatural predictions, especially for portraits with hard shadows. We overcome the limitations of existing local propagation methods by formulating the removal problem as a generation task where a diffusion model learns to globally rebuild the human appearance from scratch as a condition of an input portrait image. For robust and natural shadow removal, we propose to train the diffusion model with a compositional repurposing framework: a pre-trained text-guided image generation model is first fine-tuned to harmonize the lighting and color of the foreground with a background scene by using a background harmonization dataset; and then the model is further fine-tuned to generate a shadow-free portrait image via a shadow-paired dataset. To overcome the limitation of losing fine details in the latent diffusion model, we propose a guided-upsampling network to restore the original high-frequency details (wrinkles and dots) from the input image. To enable our compositional training framework, we construct a high-fidelity and large-scale dataset using a lightstage capturing system and synthetic graphics simulation. Our generative framework effectively removes shadows caused by both self and external occlusions while maintaining original lighting distribution and high-frequency details. Our method also demonstrates robustness to diverse subjects captured in real environments.

人像修复扩散模型去阴影高保真

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。