arXiv:2510.02423cs.AI2025-10被引 10

重构电影镜头理解评估标准,提升模型评测可靠性

RefineShot: Rethinking Cinematography Understanding with Foundational Skill Evaluation

  • 重新设计选项以消除歧义,统一评估标准
  • 发现现有模型在推理一致性上存在明显缺陷
  • 新增多维度评估协议,适合影视与多模态研究者

电影镜头理解指识别场景视觉内容及塑造叙事意义的拍摄技巧。该能力在真实世界多模态应用和影视内容创作中日益重要。当前最全面的基准ShotBench涵盖广泛电影概念并采用VQA式评估,其中ShotVL在该基准上达到领先性能。然而,我们分析发现ShotBench存在选项设计模糊问题,且ShotVL在推理一致性与指令遵循方面表现不足,导致评估不可靠,限制了公平比较与后续进展。为此,我们通过一致化选项重构系统性改进ShotBench,首次对ShotVL的推理行为进行批判性分析,并引入联合评估任务准确率与核心能力的扩展协议。由此形成RefineShot,一个更可靠、更全面的基准,推动电影镜头理解领域进一步发展。

原文摘要 · Abstract (English)

Cinematography understanding refers to the ability to recognize not only the visual content of a scene but also the cinematic techniques that shape narrative meaning. This capability is attracting increasing attention, as it enhances multimodal understanding in real-world applications and underpins coherent content creation in film and media. As the most comprehensive benchmark for this task, ShotBench spans a wide range of cinematic concepts and VQA-style evaluations, with ShotVL achieving state-of-the-art results on it. However, our analysis reveals that ambiguous option design in ShotBench and ShotVL's shortcomings in reasoning consistency and instruction adherence undermine evaluation reliability, limiting fair comparison and hindering future progress. To overcome these issues, we systematically refine ShotBench through consistent option restructuring, conduct the first critical analysis of ShotVL's reasoning behavior, and introduce an extended evaluation protocol that jointly assesses task accuracy and core model competencies. These efforts lead to RefineShot, a refined and expanded benchmark that enables more reliable assessment and fosters future advances in cinematography understanding.

电影理解多模态评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。