arXiv:2501.09960cs.CV2025-01AAAI被引 2

用离散先验建模时序一致性,提升盲脸视频修复质量

Discrete Prior-based Temporal-coherent Content Prediction for Blind Face Video Restoration

  • 基于离散视觉与运动先验,生成高质量面部内容
  • 通过帧间统计调制,使预测内容时间上更连贯
  • 适合复杂退化场景下的真实人脸视频修复

盲脸视频修复旨在从经历复杂未知退化的视频中恢复高保真细节。该任务需同时应对时序异质性并保持面部属性稳定。本文提出离散先验引导的时序一致内容预测变换器(DP-TempCoh)。设计时空感知的内容预测模块,基于退化视频令牌从离散视觉先验中合成高质量内容;进一步引入运动统计调制模块,依据跨帧均值与方差的离散运动先验调节内容,使预测结果的统计特性随时间匹配真实视频。大量实验验证了各设计的有效性,DP-TempCoh在合成与自然退化视频修复中均表现优越。

原文摘要 · Abstract (English)

Blind face video restoration aims to restore high-fidelity details from videos subjected to complex and unknown degradations. This task poses a significant challenge of managing temporal heterogeneity while at the same time maintaining stable face attributes. In this paper, we introduce a Discrete Prior-based Temporal-Coherent content prediction transformer to address the challenge, and our model is referred to as DP-TempCoh. Specifically, we incorporate a spatial-temporal-aware content prediction module to synthesize high-quality content from discrete visual priors, conditioned on degraded video tokens. To further enhance the temporal coherence of the predicted content, a motion statistics modulation module is designed to adjust the content, based on discrete motion priors in terms of cross-frame mean and variance. As a result, the statistics of the predicted content can match with that of real videos over time. By performing extensive experiments, we verify the effectiveness of the design elements and demonstrate the superior performance of our DP-TempCoh in both synthetically and naturally degraded video restoration.

视频修复时序一致人脸重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。