首个在同视频中测试多时长理解能力的基准,推动长视频模型评估革新。
ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding
- 在同一个视频中设计四层级时长问题,实现跨尺度直接比较。
- 23个多模态大模型测试显示中间时长理解能力最弱,呈倒U型表现。
- 适合研究长视频理解、多模态模型时序推理的学者使用。
长视频理解需捕捉从片段(秒级)、镜头(数十秒)、事件(分钟级)到故事(小时级)的分层时间信息。现有基准或忽略多时标设计,或在不同视频间分散时标问题,无法在同一内容上直接比较模型性能。为此,我们提出 ScaleLong,首个将四个层级时标(片段、镜头、事件、故事)的问题全部嵌入同一视频内容的基准。该设计使模型在相同视频上的跨时标表现可直接对比。ScaleLong 包含 269 个长视频(平均 86 分钟),涵盖 5 大类与 36 小类,每视频有 4–8 道精心设计的问题,至少覆盖每个时标。对 23 个 MLLMs 的评估发现,性能呈倒 U 型曲线:短时标与长时标准确率更高,中等时标最低。消融实验表明,增加视觉标记容量能持续提升各时标下的推理能力。该数据集与代码已开源。
原文摘要 · Abstract (English)
Although long-video understanding demands that models capture hierarchical temporal information -- from clip (seconds) and shot (tens of seconds) to event (minutes) and story (hours) -- existing benchmarks either neglect this multi-scale design or scatter scale-specific questions across different videos, preventing direct comparison of model performance across timescales on the same content. To address this, we introduce ScaleLong, the first benchmark to disentangle these factors by embedding questions targeting four hierarchical timescales -- clip (seconds), shot (tens of seconds), event (minutes), and story (hours) -- all within the same video content. This within-content multi-timescale questioning design enables direct comparison of model performance across timescales on identical videos. ScaleLong features 269 long videos (avg.\ 86\,min) from 5 main categories and 36 sub-categories, with 4--8 carefully designed questions, including at least one question for each timescale. Evaluating 23 MLLMs reveals a U-shaped performance curve, with higher accuracy at the shortest and longest timescales and a dip at intermediate levels. Furthermore, ablation studies show that increased visual token capacity consistently enhances reasoning across all timescales. ScaleLong offers a fine-grained, multi-timescale benchmark for advancing MLLM capabilities in long-video understanding. The code and dataset are available https://github.com/multimodal-art-projection/ScaleLong.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。