arXiv:2505.15145cs.CV2025-05NeurIPS被引 16

首个专业标注的电影技法评测基准,评估模型理解与生成影视镜头能力

CineTechBench: A Benchmark for Cinematographic Technique Understanding and Generation

  • 基于专家标注,覆盖7类电影技法的图文数据集
  • 15+大模型在理解任务中平均准确率不足50%
  • 适合影视生成、多模态理解研究者使用

电影摄像是影视制作与欣赏的核心,通过镜头运动、构图、光影等视觉元素塑造情绪与叙事。尽管多模态大模型和视频生成模型取得进展,当前模型对电影技法的理解与再现能力仍不明确,主要受限于专家标注数据稀缺。为此,我们提出CineTechBench,一个由资深电影摄影师精准标注的开创性评测基准,涵盖7个关键维度:镜头尺度、视角、构图、镜头运动、灯光、色彩与焦距。该基准包含600+幅标注电影图像和120段带明确电影技法的视频片段。针对理解任务,设计问答对与描述标注以评估多模态大模型(MLLMs)解释电影技法的能力;针对生成任务,评估先进视频生成模型在给定文本提示或关键帧条件下重建电影级镜头运动的能力。我们在15+ MLLMs和5+视频生成模型上开展大规模评估,结果揭示了现有模型的局限,并指明自动影视创作与鉴赏的未来方向。代码与数据集已开源。

原文摘要 · Abstract (English)

Cinematography is a cornerstone of film production and appreciation, shaping mood, emotion, and narrative through visual elements such as camera movement, shot composition, and lighting. Despite recent progress in multimodal large language models (MLLMs) and video generation models, the capacity of current models to grasp and reproduce cinematographic techniques remains largely uncharted, hindered by the scarcity of expert-annotated data. To bridge this gap, we present CineTechBench, a pioneering benchmark founded on precise, manual annotation by seasoned cinematography experts across key cinematography dimensions. Our benchmark covers seven essential aspects-shot scale, shot angle, composition, camera movement, lighting, color, and focal length-and includes over 600 annotated movie images and 120 movie clips with clear cinematographic techniques. For the understanding task, we design question answer pairs and annotated descriptions to assess MLLMs' ability to interpret and explain cinematographic techniques. For the generation task, we assess advanced video generation models on their capacity to reconstruct cinema-quality camera movements given conditions such as textual prompts or keyframes. We conduct a large-scale evaluation on 15+ MLLMs and 5+ video generation models. Our results offer insights into the limitations of current models and future directions for cinematography understanding and generation in automatically film production and appreciation. The code and benchmark can be accessed at https://github.com/PRIS-CV/CineTechBench.

电影技法多模态评测视频生成视觉理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。