用参考图保持人脸视频还原的身份一致性
Show and Polish: Reference-Guided Identity Preservation in Face Video Restoration
- 引入参考图像作为视觉提示,通过解耦交叉注意力保留身份细节
- 在24帧内有效减少身份漂移,跨片段采用指数融合确保连贯性
- 适合需要高保真人脸还原的影视修复、安防监控等场景
人脸视频恢复(FVR)旨在从退化版本中恢复高质量人脸视频。传统方法在严重退化时难以保持精细的身份特征,常生成缺乏个体特征的平均化人脸。为此,我们提出IP-FVR,利用高质量参考人脸图像作为视觉提示,在去噪过程中提供身份条件。IP-FVR通过解耦交叉注意力机制提取参考图像中的语义丰富身份信息,确保结果细节与身份一致。针对24帧内的帧内身份漂移,引入基于余弦相似度奖励信号与后缀加权时间聚合的身份保持反馈学习方法,有效抑制序列内漂移。对于跨片段身份漂移,设计指数融合策略,在去噪过程中迭代融合前片段帧,实现片段间身份对齐。此外,采用多流负向提示增强恢复过程,引导模型关注相关面部属性,减少低质量或错误特征生成。在合成与真实数据集上的大量实验表明,IP-FVR在质量和身份保留方面均优于现有方法,展现出在人脸视频恢复中的实际应用潜力。
原文摘要 · Abstract (English)
Face Video Restoration (FVR) aims to recover high-quality face videos from degraded versions. Traditional methods struggle to preserve fine-grained, identity-specific features when degradation is severe, often producing average-looking faces that lack individual characteristics. To address these challenges, we introduce IP-FVR, a novel method that leverages a high-quality reference face image as a visual prompt to provide identity conditioning during the denoising process. IP-FVR incorporates semantically rich identity information from the reference image using decoupled cross-attention mechanisms, ensuring detailed and identity consistent results. For intra-clip identity drift (within 24 frames), we introduce an identity-preserving feedback learning method that combines cosine similarity-based reward signals with suffix-weighted temporal aggregation. This approach effectively minimizes drift within sequences of frames. For inter-clip identity drift, we develop an exponential blending strategy that aligns identities across clips by iteratively blending frames from previous clips during the denoising process. This method ensures consistent identity representation across different clips. Additionally, we enhance the restoration process with a multi-stream negative prompt, guiding the model's attention to relevant facial attributes and minimizing the generation of low-quality or incorrect features. Extensive experiments on both synthetic and real-world datasets demonstrate that IP-FVR outperforms existing methods in both quality and identity preservation, showcasing its substantial potential for practical applications in face video restoration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。