arXiv:2504.09255cs.CV2025-04中稿 · ACM MM 2025被引 18

首个大规模人脸视频质量评估数据集及基于大模型的方法

FVQ: A Large-Scale Dataset and an LMM-based Method for Face Video Quality Assessment

  • 构建2万条真实场景人脸视频数据集,含主观评分
  • 提出FVQ-Rater模型,实现接近人眼的视频质量打分
  • 适合多媒体质量评估、AI视频审核等研究者使用

人脸视频质量评估(FVQA)因社交平台内容以人脸视频为主,且人类视觉系统对人脸敏感而具有重要意义,但受限于缺乏大规模数据集,该领域研究较少。为此,本文首次提出一个大规模真实场景下的人脸视频质量评估数据集FVQ-20K,包含20,000条真实人脸视频及其对应的平均意见得分(MOS)标注。同时,提出FVQ-Rater方法,首次探索大模态模型(LMM)在FVQA任务中的潜力。该方法提取空间、时间及人脸特有特征(如肖像特征和人脸嵌入),并采用基于LoRA的指令微调技术实现质量特定的精细化调整,在FVQ-20K与CFVQA数据集上均表现出优越性能。大量实验与分析验证了该数据集和方法在推动FVQA发展方面的巨大潜力。

原文摘要 · Abstract (English)

Face video quality assessment (FVQA) deserves to be explored in addition to general video quality assessment (VQA), as face videos are the primary content on social media platforms and human visual system (HVS) is particularly sensitive to human faces. However, FVQA is rarely explored due to the lack of large-scale FVQA datasets. To fill this gap, we present the first large-scale in-the-wild FVQA dataset, FVQ-20K, which contains 20,000 in-the-wild face videos together with corresponding mean opinion score (MOS) annotations. Along with the FVQ-20K dataset, we further propose a specialized FVQA method named FVQ-Rater to achieve human-like rating and scoring for face video, which is the first attempt to explore the potential of large multimodal models (LMMs) for the FVQA task. Concretely, we elaborately extract multi-dimensional features including spatial features, temporal features, and face-specific features (i.e., portrait features and face embeddings) to provide comprehensive visual information, and take advantage of the LoRA-based instruction tuning technique to achieve quality-specific fine-tuning, which shows superior performance on both FVQ-20K and CFVQA datasets. Extensive experiments and comprehensive analysis demonstrate the significant potential of the FVQ-20K dataset and FVQ-Rater method in promoting the development of FVQA.

视频质量人脸评估大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。