arXiv:2501.01235cs.CVcs.LG2025-01CVPR被引 7

统一框架实现视频人脸修复、补全与着色,提升时序一致性与质量。

SVFR: A Unified Framework for Generalized Video Face Restoration

  • 基于SVD生成先验,融合任务嵌入与统一潜在正则化。
  • 多任务协同使修复质量优于现有方法,时序稳定性显著提升。
  • 适合需要高质量视频人脸处理的研究与工业应用。

人脸修复(FR)是图像与视频处理中的关键领域,旨在从退化输入中重建高质量人像。尽管图像级修复进展迅速,视频级修复仍因时序一致性、运动伪影及高质量视频数据稀缺而研究不足。传统方法侧重于分辨率提升,较少关注修复、补全与着色等关联任务。本文提出广义视频人脸修复(GVFR)新任务,整合视频超分、补全与着色,并实证证明其相互促进。我们构建统一框架SVFR,利用Stable Video Diffusion(SVD)的生成与运动先验,通过统一人脸修复架构引入任务特定信息。设计可学习的任务嵌入以增强任务识别,提出新型统一潜在正则化(ULR)促进不同子任务间的共享特征学习。为提升修复质量与时序稳定性,引入人脸先验学习与自参照精修策略,用于训练与推理。该框架有效结合多任务互补优势,显著增强时序连贯性并实现更优修复效果。本工作推动视频人脸修复技术前沿,建立广义视频人脸修复新范式。代码与演示视频见https://github.com/wangzhiyaoo/SVFR.git。

原文摘要 · Abstract (English)

Face Restoration (FR) is a crucial area within image and video processing, focusing on reconstructing high-quality portraits from degraded inputs. Despite advancements in image FR, video FR remains relatively under-explored, primarily due to challenges related to temporal consistency, motion artifacts, and the limited availability of high-quality video data. Moreover, traditional face restoration typically prioritizes enhancing resolution and may not give as much consideration to related tasks such as facial colorization and inpainting. In this paper, we propose a novel approach for the Generalized Video Face Restoration (GVFR) task, which integrates video BFR, inpainting, and colorization tasks that we empirically show to benefit each other. We present a unified framework, termed as stable video face restoration (SVFR), which leverages the generative and motion priors of Stable Video Diffusion (SVD) and incorporates task-specific information through a unified face restoration framework. A learnable task embedding is introduced to enhance task identification. Meanwhile, a novel Unified Latent Regularization (ULR) is employed to encourage the shared feature representation learning among different subtasks. To further enhance the restoration quality and temporal stability, we introduce the facial prior learning and the self-referred refinement as auxiliary strategies used for both training and inference. The proposed framework effectively combines the complementary strengths of these tasks, enhancing temporal coherence and achieving superior restoration quality. This work advances the state-of-the-art in video FR and establishes a new paradigm for generalized video face restoration. Code and video demo are available at https://github.com/wangzhiyaoo/SVFR.git.

视频修复人脸修复多任务学习扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。