arXiv:2605.01496cs.CV2026-05

首个长视频问答竞赛揭示:理解故事关键在信息筛选,非模型大小。

SF20K Competition 2025: Summary and findings

论文配图:SF20K Competition 2025: Summary and findings
图 1 · 摘自论文原文
  • 用业余短片构建开放问答任务,避免记忆热门电影
  • 主赛道最佳模型准确率65.7%,特训赛道48.7%
  • 小模型多阶段设计可媲美大模型,字幕质量影响最大

本报告总结了首届短片20K(SF20K)竞赛的成果,该竞赛与ICCV 2025年SLoMO研讨会联合举办。竞赛旨在推动超越短片段动作识别的故事级视频理解,基于业余短片语料库构建开放式视频问答任务,确保模型依赖多模态理解而非记忆流行电影。评估使用SF20K-Test基准(95部影片,979个问答对),由基于GPT-4.1-nano的LLM-QA-Eval自动评分。共吸引22支队伍提交286份参赛作品,分为主赛道(无模型规模限制)和特训赛道(模型参数低于80亿)。获胜团队在主赛道取得65.7%准确率,特训赛道为48.7%,人类表现上限为91.7%。分析显示:具备叙事意识的逐镜头处理优于均匀采样;精心设计的小模型多阶段流程可匹敌或超越30倍以上参数的大模型端到端推理;字幕质量是性能主导因素。结果表明,长视频问答的核心瓶颈在于信息选择与推理结构,而非模型容量,当前方法与人类叙事理解仍存在显著差距。

原文摘要 · Abstract (English)

This report presents the results and findings of the first edition of the Short-Films 20K (SF20K) Competition, held in conjunction with the SLoMO Workshop at ICCV 2025. The competition is designed to advance story-level video understanding beyond short-clip action recognition, introducing an open-ended video question-answering task built on a corpus of amateur short films. This setup ensures that models must rely on multimodal understanding rather than memorization of popular movies. Evaluation is conducted using the SF20K-Test benchmark (95 movies, 979 question-answer pairs) and scored via LLM-QA-Eval, an automated judge based on GPT-4.1-nano. The competition attracted 22 teams and 286 submissions across two tracks: a Main Track with unrestricted model size and a Special Track limited to models under 8 billion parameters. The winning team achieved 65.7% accuracy on the Main Track and 48.7% on the Special Track, against a human performance ceiling of 91.7%. Our analysis reveals several key findings: narrative-aware, shot-level processing consistently outperforms uniform frame sampling; well-designed multi-stage pipelines using smaller models can match or exceed end-to-end inference with models over 30x larger; and subtitle quality is a dominant factor in performance. These results highlight that the primary bottleneck in long-form video QA lies in information selection and reasoning structure rather than raw model capacity, and that a substantial gap remains between current methods and human-level narrative comprehension.

视频理解问答系统多模态竞赛分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。