构建首个全面对齐人类偏好的视频生成评估基准
Video-Bench: Human-Aligned Video Generation Benchmark
- 用多模态大模型结合少样本评分与思维链,系统评估视频质量
- 在Sora等先进模型上,评估结果比传统方法更贴近人类偏好
- 适合视频生成研究者、评测人员及模型优化团队使用
视频生成评估对于确保生成模型产出视觉真实、高质量且符合人类期望的视频至关重要。现有视频生成基准分为两类:传统基准通过度量指标和嵌入向量在多个维度评估视频质量,但往往与人类判断不一致;基于大语言模型(LLM)的基准虽具备类人推理能力,却受限于对视频质量指标和跨模态一致性理解不足。为解决上述问题并建立更贴近人类偏好的评估基准,本文提出Video-Bench,一个包含丰富提示集和广泛评估维度的综合性基准。该基准首次系统性地在视频生成评估的所有相关维度上应用多模态大模型(MLLM)。通过引入少样本评分和链式查询技术,Video-Bench提供了一种结构化、可扩展的生成视频评估方法。在Sora等先进模型上的实验表明,Video-Bench在所有维度上均显著提升与人类偏好的对齐度。此外,当框架评估结果与人类评价存在差异时,其始终提供更具客观性和准确性的洞察,显示出超越传统人类判断的更大潜力。
原文摘要 · Abstract (English)
Video generation assessment is essential for ensuring that generative models produce visually realistic, high-quality videos while aligning with human expectations. Current video generation benchmarks fall into two main categories: traditional benchmarks, which use metrics and embeddings to evaluate generated video quality across multiple dimensions but often lack alignment with human judgments; and large language model (LLM)-based benchmarks, though capable of human-like reasoning, are constrained by a limited understanding of video quality metrics and cross-modal consistency. To address these challenges and establish a benchmark that better aligns with human preferences, this paper introduces Video-Bench, a comprehensive benchmark featuring a rich prompt suite and extensive evaluation dimensions. This benchmark represents the first attempt to systematically leverage MLLMs across all dimensions relevant to video generation assessment in generative models. By incorporating few-shot scoring and chain-of-query techniques, Video-Bench provides a structured, scalable approach to generated video evaluation. Experiments on advanced models including Sora demonstrate that Video-Bench achieves superior alignment with human preferences across all dimensions. Moreover, in instances where our framework's assessments diverge from human evaluations, it consistently offers more objective and accurate insights, suggesting an even greater potential advantage over traditional human judgment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。