首个评估图文视频感知与推理的基准,揭示当前模型表现严重不足。
IV-Bench: A Benchmark for Image-Grounded Video Perception and Reasoning in Multimodal LLMs
- 构建967段图文配对视频数据集,覆盖13项感知与推理任务
- 主流模型在关键任务上最高仅28.9%准确率,显著落后于人类水平
- 适合研究多模态大模型视频理解、跨模态对齐的学者使用
现有多模态大模型(MLLMs)评估框架多聚焦图像推理或通用视频理解,忽视了图像上下文在视频理解中的关键作用。为弥补这一空白,我们提出IV-Bench,首个专门评估图像锚定视频感知与推理的综合性基准。IV-Bench包含967个视频,对应2,585个精心标注的图文查询,涵盖13项任务(7项感知+6项推理)和5类典型场景。对前沿开源(如InternVL2.5、Qwen2.5-VL)与闭源模型(如GPT-4o、Gemini2-Flash、Gemini2-Pro)的广泛评测显示,当前模型在图像锚定视频任务中表现远低于预期,最高准确率仅为28.9%。进一步分析揭示推理模式、帧数、分辨率是影响性能的关键因素。此外,通过简单数据合成方法验证,IV-Bench的挑战不仅源于训练数据格式对齐问题。这些发现为未来研究提供了重要参考。代码与数据已开源:https://github.com/multimodal-art-projection/IV-Bench。
原文摘要 · Abstract (English)
Existing evaluation frameworks for Multimodal Large Language Models (MLLMs) primarily focus on image reasoning or general video understanding tasks, largely overlooking the significant role of image context in video comprehension. To bridge this gap, we propose IV-Bench, the first comprehensive benchmark for evaluating Image-Grounded Video Perception and Reasoning. IV-Bench consists of 967 videos paired with 2,585 meticulously annotated image-text queries across 13 tasks (7 perception and 6 reasoning tasks) and 5 representative categories. Extensive evaluations of state-of-the-art open-source (e.g., InternVL2.5, Qwen2.5-VL) and closed-source (e.g., GPT-4o, Gemini2-Flash and Gemini2-Pro) MLLMs demonstrate that current models substantially underperform in image-grounded video Perception and Reasoning, merely achieving at most 28.9% accuracy. Further analysis reveals key factors influencing model performance on IV-Bench, including inference pattern, frame number, and resolution. Additionally, through a simple data synthesis approach, we demonstratethe challenges of IV- Bench extend beyond merely aligning the data format in the training proecss. These findings collectively provide valuable insights for future research. Our codes and data are released in https://github.com/multimodal-art-projection/IV-Bench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。