arXiv:2602.23969cs.MMcs.CV2026-02ACL被引 10

首个面向多镜头视频生成的综合性评估基准,突破单镜头局限。

MSVBench: Towards Human-Level Evaluation of Multi-Shot Video Generation

  • 融合大模型语义理解与专用模型感知判断,实现多层级评估
  • 20个模型测试显示,当前模型仍以视觉插值为主,缺乏世界建模能力
  • 可生成可复用监督信号,轻量模型微调后媲美商业级模型

视频生成向复杂多镜头叙事发展,现有评估方法仍停留在单镜头范式,缺乏对长篇连贯性与吸引力的全面衡量。为此,我们提出MSVBench,首个针对多镜头视频生成的综合性基准,包含分层脚本和参考图像。我们设计混合评估框架,结合大模型的高层语义推理与领域专家模型的精细感知判断。在20种不同范式的视频生成方法上进行评估,发现当前模型虽具强视觉保真度,但主要表现为视觉插值器而非真实世界模型。我们进一步验证了基准可靠性:与人类判断的斯皮尔曼等级相关系数达94.4%。此外,MSVBench还提供可扩展的监督信号,对轻量模型在流水线优化后的推理轨迹进行微调,即可达到与Gemini-2.5-Flash等商用模型相当的人类对齐性能。

原文摘要 · Abstract (English)

The evolution of video generation toward complex, multi-shot narratives has exposed a critical deficit in current evaluation methods. Existing benchmarks remain anchored to single-shot paradigms, lacking the comprehensive story assets and cross-shot metrics required to assess long-form coherence and appeal. To bridge this gap, we introduce MSVBench, the first comprehensive benchmark featuring hierarchical scripts and reference images tailored for Multi-Shot Video generation. We propose a hybrid evaluation framework that synergizes the high-level semantic reasoning of Large Multimodal Models (LMMs) with the fine-grained perceptual rigor of domain-specific expert models. Evaluating 20 video generation methods across diverse paradigms, we find that current models--despite strong visual fidelity--primarily behave as visual interpolators rather than true world models. We further validate the reliability of our benchmark by demonstrating a state-of-the-art Spearman's rank correlation of 94.4% with human judgments. Finally, MSVBench extends beyond evaluation by providing a scalable supervisory signal. Fine-tuning a lightweight model on its pipeline-refined reasoning traces yields human-aligned performance comparable to commercial models like Gemini-2.5-Flash.

视频生成评估基准多镜头大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。