arXiv:2506.21356cs.CV2025-06NeurIPS被引 22

构建电影镜头理解新基准,揭示大模型在影视语言上的短板

ShotBench: Expert-Level Cinematic Understanding in Vision-Language Models

  • 设计3500+专家标注的电影镜头问答对,覆盖8大影视维度
  • 24个主流模型平均准确率不足60%,尤其在空间推理上表现差
  • 开源7万条数据与新模型,助力影视内容生成与理解

电影摄影是电影叙事、情感与美学表达的核心视觉语言。尽管当前视觉语言模型(VLMs)具备较强的通用视觉理解能力,但其对单个镜头中蕴含的细微影视语法的理解仍缺乏系统评估。为此,我们提出ShotBench,一个专用于电影语言理解的综合性基准,包含超过3.5k条来自200多部知名影片(主要为奥斯卡提名作品)的专家标注问答对,涵盖8个关键影视维度。对24个主流VLMs在ShotBench上的评估显示,即使表现最好的模型平均准确率也低于60%,尤其在细粒度视觉线索和复杂空间推理方面存在显著不足。为推动该领域发展,我们构建了包含约7万条电影问答对的ShotQA大规模多模态数据集,并基于此通过监督微调与组相对策略优化训练出ShotVL模型。ShotVL在ShotBench上显著超越所有现有开源与闭源模型,达到新的最佳性能。我们已开源模型、数据与代码,以促进人工智能在电影理解与生成领域的快速进步。

原文摘要 · Abstract (English)

Cinematography, the fundamental visual language of film, is essential for conveying narrative, emotion, and aesthetic quality. While recent Vision-Language Models (VLMs) demonstrate strong general visual understanding, their proficiency in comprehending the nuanced cinematic grammar embedded within individual shots remains largely unexplored and lacks robust evaluation. This critical gap limits both fine-grained visual comprehension and the precision of AI-assisted video generation. To address this, we introduce ShotBench, a comprehensive benchmark specifically designed for cinematic language understanding. It features over 3.5k expert-annotated QA pairs from images and video clips, meticulously curated from over 200 acclaimed (predominantly Oscar-nominated) films and spanning eight key cinematography dimensions. Our evaluation of 24 leading VLMs on ShotBench reveals their substantial limitations: even the top-performing model achieves less than 60% average accuracy, particularly struggling with fine-grained visual cues and complex spatial reasoning. To catalyze advancement in this domain, we construct ShotQA, a large-scale multimodal dataset comprising approximately 70k cinematic QA pairs. Leveraging ShotQA, we develop ShotVL through supervised fine-tuning and Group Relative Policy Optimization. ShotVL significantly outperforms all existing open-source and proprietary models on ShotBench, establishing new state-of-the-art performance. We open-source our models, data, and code to foster rapid progress in this crucial area of AI-driven cinematic understanding and generation.

电影理解视觉语言模型多模态评测数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。