构建首个覆盖长文档理解、推理与定位的综合评测基准
LongDocURL: a Comprehensive Multimodal Long Document Benchmark Integrating Understanding, Reasoning, and Locating
- 定义三大任务类别,整合20个子任务,覆盖复杂文档场景
- 收集2325组问答对,涵盖超3.3万页文档,规模远超现有数据集
- 适用于评估大模型在长文档中的综合能力,适合研究者与开发者使用
大型视觉语言模型(LVLMs)显著提升了文档理解能力,可处理复杂文档元素、长上下文和多样化任务。然而,现有评测基准仅限于少量页面,且缺乏对版面元素定位的全面分析。本文首先定义三大核心任务:长文档理解、数值推理和跨元素定位,并提出综合性基准LongDocURL,包含20个子任务,按任务类型与答案证据分类。通过半自动化构建流程,收集2,325组高质量问答对,覆盖超过33,000页文档,显著超越现有基准。进一步在26种不同配置下对开源与闭源模型进行系统评估,揭示该领域存在的关键性能差距。
原文摘要 · Abstract (English)
Large vision language models (LVLMs) have improved the document understanding capabilities remarkably, enabling the handling of complex document elements, longer contexts, and a wider range of tasks. However, existing document understanding benchmarks have been limited to handling only a small number of pages and fail to provide a comprehensive analysis of layout elements locating. In this paper, we first define three primary task categories: Long Document Understanding, numerical Reasoning, and cross-element Locating, and then propose a comprehensive benchmark, LongDocURL, integrating above three primary tasks and comprising 20 sub-tasks categorized based on different primary tasks and answer evidences. Furthermore, we develop a semi-automated construction pipeline and collect 2,325 high-quality question-answering pairs, covering more than 33,000 pages of documents, significantly outperforming existing benchmarks. Subsequently, we conduct comprehensive evaluation experiments on both open-source and closed-source models across 26 different configurations, revealing critical performance gaps in this field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。