arXiv:2604.14560cs.CV2026-04

一拍即合:用时空双先验提升视频人脸修复质量

DVFace: Spatio-Temporal Dual-Prior Diffusion for Video Face Restoration

论文配图:DVFace: Spatio-Temporal Dual-Prior Diffusion for Video Face Restoration
图 1 · 摘自论文原文
  • 设计时空双编码器,分别提取面部空间与时间先验
  • 单步扩散+不对称融合,修复速度更快且更稳定
  • 适合追求高清人脸修复与身份一致性的研究者

视频人脸修复旨在将退化的面部视频还原为高质量、细节真实、身份稳定且时间连贯的成果。近期基于扩散的方法虽引入强大生成先验,提升了细节合成能力,但仍依赖通用扩散先验和多步采样,限制了面部适应性与推理效率。为此,本文提出DVFace,一种面向真实世界视频人脸修复的一步扩散框架。通过设计时空双编码器,从退化视频中提取互补的空间与时间面部先验,并引入非对称时空融合模块,按其不同作用注入扩散主干。在多个基准上的评估表明,相比现有方法,DVFace在修复质量、时间一致性与身份保真度方面均表现更优。

原文摘要 · Abstract (English)

Video face restoration aims to enhance degraded face videos into high-quality results with realistic facial details, stable identity, and temporal coherence. Recent diffusion-based methods have brought strong generative priors to restoration and enabled more realistic detail synthesis. However, existing approaches for face videos still rely heavily on generic diffusion priors and multi-step sampling, which limit both facial adaptation and inference efficiency. These limitations motivate the use of one-step diffusion for video face restoration, yet achieving faithful facial recovery alongside temporally stable outputs remains challenging. In this paper, we propose, DVFace, a one-step diffusion framework for real-world video face restoration. Specifically, we introduce a spatio-temporal dual-codebook design to extract complementary spatial and temporal facial priors from degraded videos. We further propose an asymmetric spatio-temporal fusion module to inject these priors into the diffusion backbone according to their distinct roles. Evaluation on various benchmarks shows that DVFace delivers superior restoration quality, temporal consistency, and identity preservation compared to recent methods. Code: https://github.com/zhengchen1999/DVFace.

视频修复扩散模型人脸重建单步生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。