arXiv:2508.19026cs.CLcs.AI2025-08EMNLP被引 3

构建电影认知理解新数据集,推动AI从看懂到读懂电影。

MovieCORE: COgnitive REasoning in Movies

  • 用多个大模型协作生成需深度思考的问答对
  • 提出新评估框架,验证模型在复杂问题上提升25%
  • 适合研究电影理解、认知推理的AI学者

本文提出MovieCORE,一个面向电影内容深度认知理解的视频问答数据集。与以往侧重表层理解的数据集不同,MovieCORE聚焦需要系统2思维的问题,且严格基于视频内容。我们采用多大模型协同的智能体式头脑风暴方法,生成并优化高质量问答对。为评估数据质量,设计了涵盖深度、启发性与句法复杂度的认知测试。同时提出全面的VQA模型评估方案,用于衡量模型在深层认知任务中的表现。针对现有视频语言模型的不足,引入后训练增强模块Agentic Choice Enhancement(ACE),使模型推理能力最高提升25%。本工作推动了AI对电影内容的理解能力,并揭示了当前模型在应对复杂、细微问题时的局限性。项目页面、数据集及代码详见https://joslefaure.github.io/assets/html/moviecore.html。

原文摘要 · Abstract (English)

This paper introduces MovieCORE, a novel video question answering (VQA) dataset designed to probe deeper cognitive understanding of movie content. Unlike existing datasets that focus on surface-level comprehension, MovieCORE emphasizes questions that engage System-2 thinking while remaining specific to the video material. We present an innovative agentic brainstorming approach, utilizing multiple large language models (LLMs) as thought agents to generate and refine high-quality question-answer pairs. To evaluate dataset quality, we develop a set of cognitive tests assessing depth, thought-provocation potential, and syntactic complexity. We also propose a comprehensive evaluation scheme for assessing VQA model performance on deeper cognitive tasks. To address the limitations of existing video-language models (VLMs), we introduce an agentic enhancement module, Agentic Choice Enhancement (ACE), which improves model reasoning capabilities post-training by up to 25%. Our work contributes to advancing movie understanding in AI systems and provides valuable insights into the capabilities and limitations of current VQA models when faced with more challenging, nuanced questions about cinematic content. Our project page, dataset and code can be found at https://joslefaure.github.io/assets/html/moviecore.html.

电影理解认知推理VQA

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。