arXiv:2509.23584cs.CV2025-09被引 9

一拍即合:用单步扩散快速提升视频人脸画质

VividFace: High-Quality and Efficient One-Step Diffusion For Video Face Enhancement

  • 将多步去噪改为单步流匹配,推理速度大幅提速
  • 通过联合潜空间与像素空间的面部引导,还原细腻纹理
  • 自动生成高质量人脸数据集,提升模型泛化能力

视频人脸增强(VFE)旨在从退化视频中恢复高质量人脸,应用广泛。现有方法在视频超分辨率与生成框架下仍面临三大挑战:(1)扩散模型迭代去噪导致计算效率低下;(2)难以同时精准建模复杂面部纹理与保持时序一致性;(3)因缺乏高质量人脸视频训练数据而泛化能力弱。为此,我们提出VividFace,一种基于预训练WANX视频生成模型的一步式高效扩散框架。该框架将传统多步扩散重构成单步流匹配,直接从低质输入映射至高质量输出,显著降低推理时间。为提升面部细节恢复,提出联合潜空间-像素空间的面部聚焦训练策略,构建时空对齐的面部掩码,引导优化聚焦关键面部区域。此外,开发基于MLLM的自动化过滤流程,构建出高质量人脸视频数据集MLLM-Face90,确保模型学习真实感面部纹理。大量实验表明,VividFace在感知质量、身份保真度和时序一致性方面均优于现有方法,涵盖合成与真实世界基准。代码、模型与数据集将公开发布,以支持后续研究。

原文摘要 · Abstract (English)

Video Face Enhancement (VFE) aims to restore high-quality facial regions from degraded video sequences, enabling a wide range of practical applications. Despite substantial progress in the field, current methods that primarily rely on video super-resolution and generative frameworks continue to face three fundamental challenges: (1) computational inefficiency caused by iterative multi-step denoising in diffusion models; (2) faithfully modeling intricate facial textures while preserving temporal consistency; and (3) limited model generalization due to the lack of high-quality face video training data. To address these challenges, we propose VividFace, a novel and efficient one-step diffusion framework for VFE. Built upon the pretrained WANX video generation model, VividFace reformulates the traditional multi-step diffusion process as a single-step flow matching paradigm that directly maps degraded inputs to high-quality outputs with significantly reduced inference time. To enhance facial detail recovery, we introduce a Joint Latent-Pixel Face-Focused Training strategy that constructs spatiotemporally aligned facial masks to guide optimization toward critical facial regions in both latent and pixel spaces. Furthermore, we develop an MLLM-driven automated filtering pipeline that produces MLLM-Face90, a meticulously curated high-quality face video dataset, ensuring models learn from photorealistic facial textures. Extensive experiments demonstrate that VividFace achieves superior performance in perceptual quality, identity preservation, and temporal consistency across both synthetic and real-world benchmarks. We will publicly release our code, models, and dataset to support future research.

视频增强扩散模型人脸生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。