针对AI生成视频的视觉质量评估难题,提出新基准与评估模型。
AIGV-Assessor: Benchmarking and Evaluating the Perceptual Quality of Text-to-Video Generation with LMM

- 构建大规模数据集AIGVQA-DB,涵盖3.6万条AI视频及37万条专家评分。
- 引入AIGV-Assessor模型,基于时空特征和多模态框架精准预测视频质量。
- 适用于评估AI视频生成质量,尤其适合研究者与开发者使用。
大型多模态模型(LMM)的快速发展推动了人工智能生成视频(AIGVs)的爆发式增长,亟需专用于AIGVs的视频质量评估(VQA)模型。现有VQA模型因无法有效识别不真实物体、异常运动或视觉不一致等独特失真,在评估AIGVs感知质量方面表现不足。为此,本文首先构建了AIGVQA-DB,一个包含36,576个由15个先进文本到视频模型生成、基于1,048个多样化提示词的AIGVs的大规模数据集,并通过系统化标注流程(评分与排序)收集了至今累计37万条专家评分。基于此数据集,进一步提出AIGV-Assessor,一种利用时空特征与LMM框架的新型VQA模型,可精准捕捉AIGVs的复杂质量属性,实现高质量评分与视频对偏好预测。在AIGVQA-DB及现有AIGV数据库上的全面实验表明,AIGV-Assessor在多个感知质量维度上显著优于现有方法,达到当前最优性能。
原文摘要 · Abstract (English)
The rapid advancement of large multimodal models (LMMs) has led to the rapid expansion of artificial intelligence generated videos (AIGVs), which highlights the pressing need for effective video quality assessment (VQA) models designed specifically for AIGVs. Current VQA models generally fall short in accurately assessing the perceptual quality of AIGVs due to the presence of unique distortions, such as unrealistic objects, unnatural movements, or inconsistent visual elements. To address this challenge, we first present AIGVQA-DB, a large-scale dataset comprising 36,576 AIGVs generated by 15 advanced text-to-video models using 1,048 diverse prompts. With these AIGVs, a systematic annotation pipeline including scoring and ranking processes is devised, which collects 370k expert ratings to date. Based on AIGVQA-DB, we further introduce AIGV-Assessor, a novel VQA model that leverages spatiotemporal features and LMM frameworks to capture the intricate quality attributes of AIGVs, thereby accurately predicting precise video quality scores and video pair preferences. Through comprehensive experiments on both AIGVQA-DB and existing AIGV databases, AIGV-Assessor demonstrates state-of-the-art performance, significantly surpassing existing scoring or evaluation methods in terms of multiple perceptual quality dimensions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。