arXiv:2607.16742cs.CV2026-07中稿 · publication in IEE…被引 4

构建首个多维评估数据集并提出统一评分模型,全面衡量AI生成人像视频质量。

Multi-Dimensional Quality Assessment for AI-Generated Human-Centric Videos: Dataset and Model

论文配图:Multi-Dimensional Quality Assessment for AI-Generated Human-Centric Videos: Dataset and Model
图 1 · 摘自论文原文
  • 基于24个文本到视频模型,构建包含20万条标注的多维评估数据集
  • 在3个维度上实现超90%的评分准确率,显著优于现有方法
  • 适合视频生成研究者与模型优化工程师快速评估生成质量

AI生成的人像视频在诸多现代应用中至关重要,但常存在质量问题与语义偏差,亟需有效的质量评估。为此,我们扩展了此前的HVEval数据集,加入成对偏好标注,形成目前最大的全维度质量评估数据集HVEval+,涵盖1000个提示、20,000个由24个文本到视频(T2V)模型生成的视频,以及60,000个平均意见得分(MOS)和60,000个偏好对,覆盖空间质量、时间质量与图文一致性三个维度,并包含20,000个特定类别问答对。同时,我们提出MoE-Rater,一种受混合专家(MoE)启发的多模态大语言模型(MLLM)方法,可统一完成多维评分、成对比较与类别问答。通过引入混合投影专家(MoPE)与混合LoRA专家(MoLE),结合三阶段训练策略,有效融合多项任务,在HVEval+与Human-AGVQA数据集上均取得优越性能。大量实验与分析表明,HVEval+与MoE-Rater在推动AI生成视频质量评估及优化方面具有巨大潜力。

原文摘要 · Abstract (English)

AI-generated human-centric videos play a crucial role in a wide range of modern applications. However, they often suffer from quality issues and semantic mismatches, underscoring the importance of effective quality assessment for such videos. To this end, we extend our previous dataset HVEval with pairwise preference annotations, resulting in HVEval+, the largest holistic quality assessment dataset for AI-generated human-centric videos, which comprises 1k prompts based on a comprehensive taxonomy, 20k videos generated by 24 text-to-video (T2V) models, and extensive human annotations, including 60k mean opinion scores (MOSs) and 60k preference pairs across 3 dimensions (i.e., spatial quality, temporal quality, and text-video correspondence), as well as 20k category-specific question-answer (Q&A) pairs. Along with the HVEval+ dataset, we further propose MoE-Rater, a Mixture-of-Experts (MoE)-inspired and multimodal large language model (MLLM)-based all-in-one method that supports multi-dimensional quality rating, multi-dimensional pairwise comparison, and category-specific question answering within a single model. Specifically, we introduce Mixture of Projector Experts (MoPE) and Mixture of LoRA Experts (MoLE), together with a three-stage training strategy consisting of task-aware pre-training, task-specific adaptation, and adaptive routing optimization, to effectively unify multiple tasks, resulting in superior performance on both HVEval+ and Human-AGVQA datasets. Extensive experiments and comprehensive analysis demonstrate the significant potential of the HVEval+ dataset and the MoE-Rater method in advancing AI-generated video quality assessment and further facilitating the evaluation and optimization of T2V models.

视频评估多维评测生成质量大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。