arXiv:2608.30068cs.CV2026-08

首个评估电影导演意图的基准,揭示大模型理解创作动机的能力严重不足

TAKE 85: Testing Audiovisual filmmaKer's intEnt across 85 Hours of Film

论文配图:TAKE 85: Testing Audiovisual filmmaKer's intEnt across 85 Hours of Film
图 1 · 摘自论文原文
  • 构建85小时398部短片的专家标注问答数据集,覆盖视听创作意图
  • 最强模型仅58分(满分100),无法准确理解镜头与声音的表达目的
  • 证明单一模态不足,需多模态协同推理,适合影视分析与AI理解研究者

影片通过灯光、色彩、构图、剪辑、对白、音乐和声音等创作选择传递意义,人类能自然感知导演意图,但现有多模态大语言模型(MLLMs)的评估几乎仅关注事件描述而非为何如此呈现。本文提出 TAKE 85,首个针对导演意图理解的基准,包含398部短片(共85小时),配备专家验证的问答对,覆盖全局与细粒度的视听意图。通过受控的模态消融实验,可系统评估多模态推理能力。在主流MLLM上的测试显示:模型虽能准确识别事件与叙事,却普遍无法推断创作决策的传达作用。结果表明,导演意图是多模态理解中被忽视的关键维度——即使最强模型也仅达58/100分,且任一输入模态均不足以独立完成任务。所有代码、问答数据与模型均已公开于 https://github.com/KaiShinozakiConefrey/Take-85

原文摘要 · Abstract (English)

Films communicate through deliberate creative choices, including lighting, color, composition, editing, dialogue, music, and sound. Humans naturally interpret these signals as directorial intent, yet current multimodal large language models (MLLMs) are evaluated almost exclusively on understanding what happens rather than why it is presented that way. We introduce TAKE 85, the first benchmark for directorial-intent understanding, comprising 398 short films (85 hours) with expert-verified question-answer pairs spanning global and fine-grained visual and audio intent. Through controlled modality ablations, TAKE 85 enables systematic evaluation of multimodal reasoning. Experiments on state-of-the-art MLLMs reveal a substantial gap between perceptual recognition and intentional understanding: while models accurately describe events and narratives, they consistently fail to infer the communicative role of filmmaking decisions. Our results establish directorial intent as a previously overlooked dimension of multimodal understanding: even the strongest model reaches only 58 out of 100, and our ablations show that no input modality is sufficient on its own. All code, Q&As, and models are publicly available from https://github.com/KaiShinozakiConefrey/Take-85

多模态理解电影分析意图识别评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。