首个全面评估长上下文视觉语言模型的基准,覆盖13000+任务实例。
MMLongBench: Benchmarking Long-Context Vision-Language Models Effectively and Thoroughly
- 构建跨模态分词方案,统一图像与文本输入长度至8K-128K token
- 46个模型测试显示:单任务表现无法代表整体长上下文能力
- 揭示推理能力越强的模型,长上下文表现越好,适合研究者与开发者参考
大视觉语言模型上下文窗口的快速扩展催生了长上下文视觉语言模型(LCVLMs),可在一个前向传播中处理数百张图像与交错文本。本文提出MMLongBench,首个覆盖多样化长上下文视觉语言任务的基准,包含13,331个样本,涵盖五类下游任务,如视觉RAG和多示例提示(Many-Shot ICL)。其图像类型包括自然与合成图像,通过跨模态分词方案将所有样本以5种标准化输入长度(8K–128K token)呈现。对46个闭源与开源LCVLM的全面评测表明:i)单任务性能是整体长上下文能力的弱代理;ii)闭源与开源模型均在长上下文任务中面临挑战,仍有巨大提升空间;iii)具备更强推理能力的模型表现出更优的长上下文性能。通过广泛的任务覆盖、多样图像类型与严格的长度控制,MMLongBench为诊断与推动下一代LCVLM发展提供了关键基础。
原文摘要 · Abstract (English)
The rapid extension of context windows in large vision-language models has given rise to long-context vision-language models (LCVLMs), which are capable of handling hundreds of images with interleaved text tokens in a single forward pass. In this work, we introduce MMLongBench, the first benchmark covering a diverse set of long-context vision-language tasks, to evaluate LCVLMs effectively and thoroughly. MMLongBench is composed of 13,331 examples spanning five different categories of downstream tasks, such as Visual RAG and Many-Shot ICL. It also provides broad coverage of image types, including various natural and synthetic images. To assess the robustness of the models to different input lengths, all examples are delivered at five standardized input lengths (8K-128K tokens) via a cross-modal tokenization scheme that combines vision patches and text tokens. Through a thorough benchmarking of 46 closed-source and open-source LCVLMs, we provide a comprehensive analysis of the current models' vision-language long-context ability. Our results show that: i) performance on a single task is a weak proxy for overall long-context capability; ii) both closed-source and open-source models face challenges in long-context vision-language tasks, indicating substantial room for future improvement; iii) models with stronger reasoning ability tend to exhibit better long-context performance. By offering wide task coverage, various image types, and rigorous length control, MMLongBench provides the missing foundation for diagnosing and advancing the next generation of LCVLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。