测试大模型读长文档能力,用图文混合难题挑战视觉语言模型
Document Haystack: A Long Context Multimodal Image/Document Understanding Vision LLM Benchmark
- 设计含5至200页的复杂文档,嵌入图文混合难点
- 8250道题评估模型在长文本中精准定位信息的能力
- 适合研究长文档理解、多模态推理的学者和开发者
多模态大语言模型显著提升了对跨模态复杂数据的分析能力,但长文档处理仍缺乏系统评估。为此,我们提出Document Haystack,一个全面的基准测试集,用于评估视觉语言模型(VLM)在长而复杂的文档上的表现。该基准包含5至200页的文档,通过在不同深度插入纯文本或图文混合的“针”(needles)来考验模型的检索能力。共涵盖400种文档变体和8,250个问题,支持客观、自动化的评估框架。本文详述了数据集构建过程与特性,并报告了主流VLM的表现,探讨了未来研究方向。
原文摘要 · Abstract (English)
The proliferation of multimodal Large Language Models has significantly advanced the ability to analyze and understand complex data inputs from different modalities. However, the processing of long documents remains under-explored, largely due to a lack of suitable benchmarks. To address this, we introduce Document Haystack, a comprehensive benchmark designed to evaluate the performance of Vision Language Models (VLMs) on long, visually complex documents. Document Haystack features documents ranging from 5 to 200 pages and strategically inserts pure text or multimodal text+image "needles" at various depths within the documents to challenge VLMs' retrieval capabilities. Comprising 400 document variants and a total of 8,250 questions, it is supported by an objective, automated evaluation framework. We detail the construction and characteristics of the Document Haystack dataset, present results from prominent VLMs and discuss potential research avenues in this area.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。