用频域感知与残差迭代,提升低质视频人脸修复的清晰度和连贯性。
FADRA: Frequency-Aware Diffusion with Residual Adaptation for Video Face Restoration

- 通过轻量LoRA+像素对齐融合,适配预训练扩散模型到人脸修复任务。
- 引入逐步残差精修头,重复利用低质输入恢复细节并保持时间一致性。
- 频域感知损失强化敏感频率成分,减少视觉抖动,适合影视修复场景。
视频人脸修复(VFR)旨在从严重退化的视频序列中恢复高质量且时间一致的面部细节;然而,现有方法在复杂退化下仍难以兼顾空间保真度与时间连贯性。为此,本文提出FADRA,一种面向鲁棒VFR的频域感知扩散框架,结合迭代残差自适应。首先,利用预训练文本到视频扩散模型的时间一致性,引入轻量级LoRA适配器与低质量(LQ)像素对齐特征融合模块,高效适配冻结生成先验。为进一步超越基于LoRA的适配,提出重复残差自适应头(RRAH),在扩散主干后进行分步残差优化。RRAH以当前速度预测与LQ隐状态为输入,使模型在每一步流匹配中反复回溯低质线索,预测残差更新,从而恢复精细面部细节并保留预训练模型的时间先验。此外,为保障重要视觉细节的结构完整性,设计频域感知损失,在多频带提供显式监督,强调对感知质量关键且易产生时间抖动的敏感频率成分。大量实验表明,FADRA在定量指标与视觉感知上均优于现有最先进方法,显著提升面部结构恢复能力与视频时间一致性。
原文摘要 · Abstract (English)
Video face restoration (VFR) aims to recover high-quality and temporally consistent facial details from severely degraded video sequences; however, existing methods still struggle to balance spatial fidelity and temporal coherence under complex degradations. To address this, we propose FADRA, a frequency-aware diffusion framework with iterative residual adaptation specifically tailored for robust VFR. We first leverage the strong temporal consistency of a pre-trained text-to-video diffusion model and introduce lightweight LoRA adapters together with a Low-Quality (LQ) Pixel-Alignment Feature Fusion module to efficiently adapt the frozen generative prior to the VFR task. To further adapt the frozen diffusion backbone to the downstream VFR task beyond LoRA-based adaptation, we introduce a Repeated Residual Adaptation Head (RRAH) for step-wise residual refinement after the diffusion backbone. To make this refinement explicitly guided by the degraded observation, RRAH further takes the LQ latent together with the current velocity prediction as input, allowing the model to repeatedly revisit LQ cues and predict residual updates at each flow-matching step. This LQ-guided repeated residual adaptation helps recover fine facial details while preserving the inherent temporal priors of the pre-trained model. Furthermore, to ensure the structural integrity of perceptually important details, we introduce a Frequency-Aware Loss that provides explicit supervision across multiple spectral bands, emphasizing visually sensitive frequency components that are crucial for perceptual quality and prone to temporal jittering. Extensive experiments demonstrate that FADRA recovers better facial structures and produces more temporally consistent videos than state-of-the-art methods, leading to clear gains in both quantitative metrics and visual perception.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。