构建首个覆盖9大任务的长视频理解基准,评估多模态大模型长期序列建模能力。
ALLVB: All-in-One Long Video Understanding Benchmark
- 将9类视频理解任务统一转为视频问答格式,实现一库多评
- 自动生成25.2万条问答,平均视频时长达近2小时,规模居首
- 适合评估长视频理解模型,推动多模态大模型发展
从图像到视频理解,多模态大模型(MLLMs)的能力日益强大。然而,现有视频理解基准普遍较短,难以有效评估MLLMs在长序列建模方面的能力。为此,我们提出ALLVB(ALL-in-One Long Video Understanding Benchmark)。其主要贡献包括:1)整合9个主流视频理解任务,并将其转化为视频问答格式,使单一基准可评估MLLMs的9种不同能力,体现其多功能性、全面性和挑战性;2)设计基于GPT-4o的全自动标注流程,仅需人工质量控制,便于持续维护与扩展;3)包含1,376个视频,覆盖16个类别,平均每段近2小时,总计25.2万条问答。据我们所知,这是目前规模最大的长视频理解基准,涵盖视频数量、平均时长和问答总量均领先。我们在ALLVB上测试了多种主流MLLMs,结果表明即使最先进的商业模型仍有巨大提升空间,凸显该基准的挑战性及长视频理解领域的巨大发展潜力。
原文摘要 · Abstract (English)
From image to video understanding, the capabilities of Multi-modal LLMs (MLLMs) are increasingly powerful. However, most existing video understanding benchmarks are relatively short, which makes them inadequate for effectively evaluating the long-sequence modeling capabilities of MLLMs. This highlights the urgent need for a comprehensive and integrated long video understanding benchmark to assess the ability of MLLMs thoroughly. To this end, we propose ALLVB (ALL-in-One Long Video Understanding Benchmark). ALLVB's main contributions include: 1) It integrates 9 major video understanding tasks. These tasks are converted into video QA formats, allowing a single benchmark to evaluate 9 different video understanding capabilities of MLLMs, highlighting the versatility, comprehensiveness, and challenging nature of ALLVB. 2) A fully automated annotation pipeline using GPT-4o is designed, requiring only human quality control, which facilitates the maintenance and expansion of the benchmark. 3) It contains 1,376 videos across 16 categories, averaging nearly 2 hours each, with a total of 252k QAs. To the best of our knowledge, it is the largest long video understanding benchmark in terms of the number of videos, average duration, and number of QAs. We have tested various mainstream MLLMs on ALLVB, and the results indicate that even the most advanced commercial models have significant room for improvement. This reflects the benchmark's challenging nature and demonstrates the substantial potential for development in long video understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。