arXiv:2410.21311cs.CVcs.AI2024-10被引 21

构建文档图文理解新基准,评估大模型细粒度视觉能力

MMDocBench: Benchmarking Large Vision-Language Models for Fine-Grained Visual Document Understanding

  • 用多粒度文档图像设计15项无OCR任务
  • 涵盖4338个问答对和11353个支持区域
  • 适合研究文档理解与多模态模型的开发者

大型视觉语言模型在众多视觉语言任务中表现优异,但在细粒度视觉理解方面仍缺乏充分评估。现有基准或包含有限且混杂的细粒度样本,或仅限于自然图像中的物体级评估。为全面评估大模型的细粒度视觉理解能力,我们提出使用包含多粒度和多模态信息的文档图像来补充自然图像。基于此,我们构建了MMDocBench,一个用于评估细粒度视觉感知与推理能力的基准,包含15个主要任务,共4338个问答对和11353个支持区域,覆盖研究论文、收据、财务报告、维基百科表格、图表和信息图等多种文档类型。基于该基准,我们对13个开源和3个专有先进视觉语言模型进行了广泛实验,评估其在不同任务和文档类型上的表现优劣。该基准、任务指令和评估代码将公开发布。

原文摘要 · Abstract (English)

Large Vision-Language Models (LVLMs) have achieved remarkable performance in many vision-language tasks, yet their capabilities in fine-grained visual understanding remain insufficiently evaluated. Existing benchmarks either contain limited fine-grained evaluation samples that are mixed with other data, or are confined to object-level assessments in natural images. To holistically assess LVLMs' fine-grained visual understanding capabilities, we propose using document images with multi-granularity and multi-modal information to supplement natural images. In this light, we construct MMDocBench, a benchmark with various OCR-free document understanding tasks for the evaluation of fine-grained visual perception and reasoning abilities. MMDocBench defines 15 main tasks with 4,338 QA pairs and 11,353 supporting regions, covering various document images such as research papers, receipts, financial reports, Wikipedia tables, charts, and infographics. Based on MMDocBench, we conduct extensive experiments using 13 open-source and 3 proprietary advanced LVLMs, assessing their strengths and weaknesses across different tasks and document image types. The benchmark, task instructions, and evaluation code will be made publicly available.

文档理解视觉语言模型评测基准细粒度分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。