提出多层级语义感知模型,精准评估AI生成视频质量
Multilevel Semantic-Aware Model for AI-Generated Video Quality Assessment
- 分帧、片段、视频三层分析,融合CLIP语义监督与交叉注意力
- 在VQAv2、AVQA等数据集上达到最佳性能,超越现有方法
- 适合关注AI视频质量评估、生成内容可信度的研究者
扩散模型的快速发展显著提升了AI生成视频的长度与连贯性,但其质量评估仍具挑战。现有方法多针对用户生成内容,极少聚焦AI生成视频的评估。本文提出MSA-VQA:一种多层级语义感知的AI生成视频质量评估模型,利用基于CLIP的语义监督与交叉注意力机制。该框架在帧、片段和视频三个层次分析内容,设计提示语义监督模块,通过CLIP文本编码器确保视频与条件提示的语义一致性;并提出语义变异感知模块,捕捉帧间细微差异。大量实验表明,该方法在VQAv2、AVQA等数据集上达到当前最优性能。
原文摘要 · Abstract (English)
The rapid development of diffusion models has greatly advanced AI-generated videos in terms of length and consistency recently, yet assessing AI-generated videos still remains challenging. Previous approaches have often focused on User-Generated Content(UGC), but few have targeted AI-Generated Video Quality Assessment methods. In this work, we introduce MSA-VQA, a Multilevel Semantic-Aware Model for AI-Generated Video Quality Assessment, which leverages CLIP-based semantic supervision and cross-attention mechanisms. Our hierarchical framework analyzes video content at three levels: frame, segment, and video. We propose a Prompt Semantic Supervision Module using text encoder of CLIP to ensure semantic consistency between videos and conditional prompts. Additionally, we propose the Semantic Mutation-aware Module to capture subtle variations between frames. Extensive experiments demonstrate our method achieves state-of-the-art results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。