arXiv:2411.09293cs.CV2024-11被引 2

用视觉语言模型提升人脸超分辨率,融合语义与深度信息

LLV-FSR: Exploiting Large Language-Vision Prior for Face Super-resolution

  • 引入预训练视觉语言模型生成图文描述、语义掩码和深度图等多模态先验
  • 在MMCelebA-HQ上实现9.17dB PSNR,比当前最优方法高0.43dB
  • 适合需要高质量人脸重建的生成与修复场景

现有面部超分辨率(FSR)方法虽有进展,但主要依赖有限视觉信息,常忽略高阶语义、深度及文本描述等多源线索,导致难以生成统一有意义的表征。本文提出LLV-FSR框架,将大规模视觉语言模型与高阶视觉先验结合,通过预训练模型生成图像描述、面部语义掩码和深度图等多模态先验,指导特征表示学习,从而实现更真实、高质量的人脸超分。实验表明,该方法在MMCelebA-HQ数据集上以9.17dB PSNR超越当前最优方案0.43dB,显著提升重建与感知质量。

原文摘要 · Abstract (English)

Existing face super-resolution (FSR) methods have made significant advancements, but they primarily super-resolve face with limited visual information, original pixel-wise space in particular, commonly overlooking the pluralistic clues, like the higher-order depth and semantics, as well as non-visual inputs (text caption and description). Consequently, these methods struggle to produce a unified and meaningful representation from the input face. We suppose that introducing the language-vision pluralistic representation into unexplored potential embedding space could enhance FSR by encoding and exploiting the complementarity across language-vision prior. This motivates us to propose a new framework called LLV-FSR, which marries the power of large vision-language model and higher-order visual prior with the challenging task of FSR. Specifically, besides directly absorbing knowledge from original input, we introduce the pre-trained vision-language model to generate pluralistic priors, involving the image caption, descriptions, face semantic mask and depths. These priors are then employed to guide the more critical feature representation, facilitating realistic and high-quality face super-resolution. Experimental results demonstrate that our proposed framework significantly improves both the reconstruction quality and perceptual quality, surpassing the SOTA by 0.43dB in terms of PSNR on the MMCelebA-HQ dataset.

人脸超分视觉语言模型多模态先验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。