VideoScore2 通过多维度分析和可解释推理,更精准评估生成视频质量。
VideoScore2: Think before You Score in Generative Video Evaluation
- 分视觉质量、文本对齐、物理一致性三维度评估,生成推理链。
- 在自建基准上准确率达44.35,跨域平均性能提升至50.37。
- 适合需要可解释性评价的生成模型优化与奖励建模场景。
文本到视频生成技术日益成熟,但视频质量评估仍面临挑战,因其涉及视觉质量、语义对齐与物理一致性等多方面。现有评估方法多为单一模糊分数,缺乏可解释性或仅提供粗略分析。我们提出 VideoScore2,一种多维度、可解释且符合人类判断的评估框架,显式评估视觉质量、文本-视频对齐及物理/常识一致性,并生成详细思维链推理。模型基于包含27,168条人工标注视频的大型数据集 VideoFeedback2 训练,采用监督微调与组相对策略优化(GRPO)的两阶段强化学习流程,提升分析鲁棒性。实验表明,VideoScore2 在自域基准 VideoScore-Bench-v2 上准确率达44.35(+5.94),在四个跨域基准(VideoGenReward-Bench、VideoPhy2等)上平均性能达50.37(+4.32),并提供可解释评估,有效支持 Best-of-N 采样中的可控生成。
原文摘要 · Abstract (English)
Recent advances in text-to-video generation have produced increasingly realistic and diverse content, yet evaluating such videos remains a fundamental challenge due to their multi-faceted nature encompassing visual quality, semantic alignment, and physical consistency. Existing evaluators and reward models are limited to single opaque scores, lack interpretability, or provide only coarse analysis, making them insufficient for capturing the comprehensive nature of video quality assessment. We present VideoScore2, a multi-dimensional, interpretable, and human-aligned framework that explicitly evaluates visual quality, text-to-video alignment, and physical/common-sense consistency while producing detailed chain-of-thought rationales. Our model is trained on a large-scale dataset VideoFeedback2 containing 27,168 human-annotated videos with both scores and reasoning traces across three dimensions, using a two-stage pipeline of supervised fine-tuning followed by reinforcement learning with Group Relative Policy Optimization (GRPO) to enhance analytical robustness. Extensive experiments demonstrate that VideoScore2 achieves superior performance with 44.35 (+5.94) accuracy on our in-domain benchmark VideoScore-Bench-v2 and 50.37 (+4.32) average performance across four out-of-domain benchmarks (VideoGenReward-Bench, VideoPhy2, etc), while providing interpretable assessments that bridge the gap between evaluation and controllable generation through effective reward modeling for Best-of-N sampling. Project Page: https://tiger-ai-lab.github.io/VideoScore2/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。