用参考图像保持人脸身份,修复模糊、压缩等退化视频。
RGFVR: Reference-Guided Face Video Restoration with Flow Matching

- 引入双模态身份条件,结合参考图与描述信息指导修复。
- 在多种退化条件下,身份保真度和时序一致性显著提升。
- 适用于需要保留特定人物特征的视频修复场景。
从退化的人脸视频中恢复高质量序列极具挑战性,需同时保证视觉清晰度、时序一致性和身份准确性。现有方法或无参考,易丢失个体面部细节;或仅针对特定人物,泛化能力差。本文提出一种不依赖具体主体的参考引导框架,将双模态感知-描述性身份条件注入预训练的流模型文本到视频生成器,并采用两阶段训练策略强化身份引导。实验表明,该方法在下采样、模糊、噪声及压缩伪影等复杂退化条件下,均显著提升重建质量、时序一致性和身份保真度,性能优于现有方法。代码已开源:https://github.com/batuhanntosun/RG-FVR。
原文摘要 · Abstract (English)
Face video restoration from degraded observations is challenging, as it requires simultaneously recovering visual fidelity, temporal consistency, and subject identity. Existing approaches are often either reference-free, which can lead to identity loss when person-specific facial details are lost, or subject-specific, which limits generalization to unseen identities. We propose a subject-agnostic, reference-guided framework for identity-preserving face video restoration. Our method introduces bimodal perceptual-descriptive identity conditioning into a pretrained flow-based text-to-video generator and employs a two-stage training strategy to strengthen identity guidance during restoration. Experiments show that our approach improves restoration fidelity, temporal consistency, and identity preservation, achieving superior performance under challenging video degradations, including downsampling, blur, noise, and compression artifacts. The code is available under: https://github.com/batuhanntosun/RG-FVR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。