arXiv:2508.04161cs.CVcs.MM2025-08AAAI

利用音视频协同提升人脸视频修复质量

Audio-Assisted Face Video Restoration with Temporal and Identity Complementary Learning

  • 分两阶段:先低分辨率时序恢复,再高分辨率身份特征精修
  • 在多个任务上超越现有方法,尤其在压缩失真、模糊和超分上表现优异
  • 适合音视频修复、媒体增强领域研究者与工程师参考

伴随音频的人脸视频在日常生活中日益普及,但常面临复杂退化问题。现有修复方法多忽视视觉与音频特征间的内在关联,尤其在嘴部区域。少数音视频协同方法仅针对压缩伪影修复。本文提出通用音视频人脸修复网络(GAVN),通过身份与时间互补学习应对多种流媒体视频退化。GAVN首先在低分辨率空间捕捉帧间时序特征以粗略恢复并节省计算开销;随后在高分辨率空间借助音频信号与人脸关键点提取帧内身份特征,实现更精细的面部细节恢复;最后重建模块融合时序与身份特征生成高质量人脸视频。实验表明,GAVN在压缩伪影去除、去模糊和超分辨率任务上均优于当前最先进方法。代码将在发表后公开。

原文摘要 · Abstract (English)

Face videos accompanied by audio have become integral to our daily lives, while they often suffer from complex degradations. Most face video restoration methods neglect the intrinsic correlations between the visual and audio features, especially in mouth regions. A few audio-aided face video restoration methods have been proposed, but they only focus on compression artifact removal. In this paper, we propose a General Audio-assisted face Video restoration Network (GAVN) to address various types of streaming video distortions via identity and temporal complementary learning. Specifically, GAVN first captures inter-frame temporal features in the low-resolution space to restore frames coarsely and save computational cost. Then, GAVN extracts intra-frame identity features in the high-resolution space with the assistance of audio signals and face landmarks to restore more facial details. Finally, the reconstruction module integrates temporal features and identity features to generate high-quality face videos. Experimental results demonstrate that GAVN outperforms the existing state-of-the-art methods on face video compression artifact removal, deblurring, and super-resolution. Codes will be released upon publication.

视频修复音视频协同人脸重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。