构建细粒度电影问答基准,测试模型对长叙事的深层理解能力
Cinéaste: A Fine-grained Contextual Movie Question Answering Benchmark
- 基于200部电影1805个场景生成3119道多选题,覆盖五类细粒度推理
- 用GPT-4o生成上下文丰富的题目,结合视觉描述与剧情摘要,要求深度理解
- 现有大模型在该数据集上最高仅63.15%准确率,暴露长时序推理瓶颈
尽管视觉语言模型在视频理解方面取得进展,但诊断其对深层叙事理解的能力仍具挑战。现有基准多测试短片段识别或使用模板化问题,难以评估长篇叙事中的细粒度推理能力。为此,我们提出$ ext{Cin}éaste$,一个面向长篇电影理解的综合性基准。数据集包含从200部多样电影中提取的1,805个场景,生成3,119道多选题,涵盖五类新颖的细粒度上下文推理类别。通过整合视觉描述、字幕、场景标题和剧情摘要,利用GPT-4o生成多样化且富含上下文的问题,要求深入的叙事理解。为确保评测质量,采用两阶段过滤流程:上下文独立性过滤确保问题依赖视频内容;上下文真实性过滤验证答案与影片内容的一致性,减少幻觉。实验表明,现有多模态大模型在$ ext{Cin}éaste$上表现不佳,分析显示长时序推理是主要瓶颈,最先进开源模型准确率仅为63.15%,凸显细粒度上下文理解的重大挑战及长篇电影理解能力提升的必要性。
原文摘要 · Abstract (English)
While recent advancements in vision-language models have improved video understanding, diagnosing their capacity for deep, narrative comprehension remains a challenge. Existing benchmarks often test short-clip recognition or use template-based questions, leaving a critical gap in evaluating fine-grained reasoning over long-form narrative content. To address these gaps, we introduce $\mathsf{Cin\acute{e}aste}$, a comprehensive benchmark for long-form movie understanding. Our dataset comprises 3,119 multiple-choice question-answer pairs derived from 1,805 scenes across 200 diverse movies, spanning five novel fine-grained contextual reasoning categories. We use GPT-4o to generate diverse, context-rich questions by integrating visual descriptions, captions, scene titles, and summaries, which require deep narrative understanding. To ensure high-quality evaluation, our pipeline incorporates a two-stage filtering process: Context-Independence filtering ensures questions require video context, while Contextual Veracity filtering validates factual consistency against the movie content, mitigating hallucinations. Experiments show that existing MLLMs struggle on $\mathsf{Cin\acute{e}aste}$; our analysis reveals that long-range temporal reasoning is a primary bottleneck, with the top open-source model achieving only 63.15\% accuracy. This underscores significant challenges in fine-grained contextual understanding and the need for advancements in long-form movie comprehension.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。