评测大模型在科学视频中的高级推理能力,发现现有模型表现不足。
SciVideoBench: Benchmarking Scientific Video Reasoning in Large Multimodal Models
- 构建1000个科学实验视频的多选题,涵盖25个专业领域。
- 顶尖模型如Gemini 2.5 Pro在复杂推理任务中表现不佳。
- 适合关注科学AI、多模态认知能力的研究者使用。
大型多模态模型(LMMs)在诸多能力上已取得显著进展,但科学领域的复杂视频推理仍是重大挑战。现有视频基准主要针对通用场景,依赖感知/识别,推理任务简单,导致评估饱和,无法有效检验高级多模态认知能力。为此,我们提出SciVideoBench,一个专为科学情境下高级视频推理设计的严谨基准。该基准包含1000个从前沿科学实验视频中精心构建的多选题,覆盖25个以上专业学科,并通过半自动系统验证。每道题需结合领域知识、精准时空感知与复杂逻辑推理,充分考验模型的高阶认知能力。评估显示,包括Gemini 2.5 Pro和Qwen2.5-VL在内的先进开源与闭源模型均存在显著性能缺陷,表明其视频推理能力仍有巨大提升空间。对推理复杂度与视觉定位等关键因素的分析,为未来LMM发展提供了清晰方向,推动真正具备科研协作能力的多模态AI诞生。我们希望SciVideoBench能激发社区兴趣,助力前沿人工智能突破科学边界。
原文摘要 · Abstract (English)
Large Multimodal Models (LMMs) have achieved remarkable progress across various capabilities; however, complex video reasoning in the scientific domain remains a significant and challenging frontier. Current video benchmarks predominantly target general scenarios where perception/recognition is heavily relied on, while with relatively simple reasoning tasks, leading to saturation and thus failing to effectively evaluate advanced multimodal cognitive skills. To address this critical gap, we introduce SciVideoBench, a rigorous benchmark specifically designed to assess advanced video reasoning in scientific contexts. SciVideoBench consists of 1,000 carefully crafted multiple-choice questions derived from cutting-edge scientific experimental videos spanning over 25 specialized academic subjects and verified by a semi-automatic system. Each question demands sophisticated domain-specific knowledge, precise spatiotemporal perception, and intricate logical reasoning, effectively challenging models' higher-order cognitive abilities. Our evaluation highlights significant performance deficits in state-of-the-art proprietary and open-source LMMs, including Gemini 2.5 Pro and Qwen2.5-VL, indicating substantial room for advancement in video reasoning capabilities. Detailed analyses of critical factors such as reasoning complexity and visual grounding provide valuable insights and clear direction for future developments in LMMs, driving the evolution of truly capable multimodal AI co-scientists. We hope SciVideoBench could fit the interests of the community and help to push the boundary of cutting-edge AI for border science.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。