首个统一评估生成视频质量与归因的基准与模型
Q-Save: Towards Scoring and Attribution for Generated Video Evaluation
- 构建包含近万条视频的标注数据集,支持三维度评分与解释
- 在三个核心维度上显著优于现有方法,且可提供可解释理由
- 适合关注生成视频质量评估与可解释性的研究者使用
评估AI生成视频(AIGV)质量涉及视觉质量、动态质量与文本-视频对齐三个关键维度。尽管已有众多评估数据集和算法,但现有方法受限于两个问题:评估维度缺乏系统定义,且三者被分别处理于独立模型中。为此,我们提出Q-Save,一个全面的基准数据集与统一评估模型。Q-Save数据集包含近10,000个视频样本,每个样本均带有平均意见分(MOS)及三个核心维度的细粒度归因解释。基于该标注数据集,我们训练了采用SlowFast框架的Q-Save模型,通过三阶段训练策略(监督微调、分组相对策略优化、最终监督微调)联合完成质量评分与归因生成,兼顾精度与效率。实验表明,Q-Save在生成视频质量预测上表现优异,并能提供可解释的推理依据。代码与数据集将在发表后公开。
原文摘要 · Abstract (English)
Evaluating AI-generated video (AIGV) quality hinges on three crucial dimensions: visual quality, dynamic quality, and text-video alignment. While numerous evaluation datasets and algorithms have been proposed, existing approaches are constrained by two limitations: the absence of systematic definitions for evaluation dimensions, and the isolated treatment of the three dimensions in separate models. Therefore, we introduce Q-Save, a holistic benchmark dataset and unified evaluation model for AIGV quality assessment. The Q-Save dataset contains nearly 10,000 video samples, each annotated with Mean Opinion Scores (MOS) and fine-grained attribution explanations across the three core dimensions. Leveraging this attribution-annotated dataset, we train the proposed Q-Save model, which adopts the SlowFast framework to balance accuracy and efficiency, and employs a three-stage training strategy with Chain-of-Thought (COT) formatted data: Supervised Fine-Tuning (SFT), Grouped Relative Policy Optimization (GRPO), and a final SFT round for stability, to jointly perform quality scoring and attribution generation. Experimental results demonstrate that Q-Save achieves superior performance in AIGV quality prediction while providing interpretable justifications. Code and dataset will be released upon publication.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。