arXiv:2608.21430cs.AIcs.CL2026-08被引 2

构建首个公开好莱坞影片叙事理解多模态评测集,揭示模型在电影叙事理解上的短板。

Evaluating Multimodal Narrative Understanding of Popular Hollywood Films

论文配图:Evaluating Multimodal Narrative Understanding of Popular Hollywood Films
图 1 · 摘自论文原文
  • 基于票房与版权状态筛选1922-1979年好莱坞影片,构建开放数据集。
  • 多模态模型在叙事理解任务中表现不佳,最高认知准确率仅61.1%。
  • 适合研究影视叙事、多模态理解与开放数据集建设的学者参考。

多模态语言模型在电影大规模计算分析方面展现出巨大潜力,为电影史研究和叙事手法演变提供了新途径。然而,受版权保护限制,围绕好莱坞电影建立稳定基准面临挑战。本文直接回应此问题,构建了一个新的好莱坞电影集合,依据两个标准:票房热度(发布首个1922-1979年《综艺》周刊周票房数据的大型开源集合);以及可能的公共领域状态(通过研究美国版权登记与续展记录确定)。在此基础上,我们构建了一个聚焦叙事元素的多模态选择题(MCQ)评测集,旨在评估模型在支持电影叙事研究方面的能力。结果显示,许多视觉-语言模型在此任务上表现堪忧,多数接近随机水平;而使用音频的视听模型最高达到61.1%的准确率,仍远低于人类表现。

原文摘要 · Abstract (English)

Multimodal language models increasingly show promise for enabling the large-scale computational analysis of film, opening up new avenues for learning about film history and the evolution of narrative techniques. But the creation of stable benchmarks built around Hollywood films is complicated by copyright protections. In this work, we address these concerns directly, by building a new collection of Hollywood films defined by two criteria: box office popularity (where we publish the first large-scale, open collection of weekly box office earnings reported by Variety magazine from 1922-1979); and likely public domain status (by researching copyright registrations and renewals in the US Catalog of Copyright Entries). We build a new multimodal MCQ benchmark on top of this collection that focuses on narrative elements that directly evaluate the abilities of models to inform meaningful research on film narrative; we find that many vision-language models struggle on this task (with many performing at near-chance levels of accuracy), while audio-visual models (including those that use audio in captioning scenes) reach a maximum accuracy of 61.1%, well below human-level performance.

多模态电影分析评测集叙事理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。