arXiv:2603.23511cs.CLcs.AI2026-03中稿 · ICLR

DISCO评估文档处理模型在多类文档上的表现,指导选型。

DISCO: Document Intelligence Suite for COmparative Evaluation

  • 分项评估OCR与视觉语言模型在文档解析和问答中的表现。
  • 手写体和长文档适合用OCR,多语言和图文混排更适合视觉语言模型。
  • 根据文档结构选择策略,避免盲目使用提示词提升性能。

文档智能需要准确的文本提取和可靠的文档内容推理。我们提出 extbf{DISCO}(Document Intelligence Suite for COmparative Evaluation),专门用于在多样化文档类型上独立评估光学字符识别(OCR)流程和视觉语言模型(VLMs)在文档解析与问答任务中的表现,涵盖手写文本、多语言文字、医疗表格、信息图以及多页文档。评估结果显示,不同任务与文档特征下性能差异显著,凸显了根据复杂度选择方法的重要性。总体而言,OCR在手写体及长篇或多页文档中表现更可靠,因显式文本定位支持以文本为中心的推理;而视觉语言模型在多语言文本和视觉丰富的布局中表现更优。任务感知提示对部分文档类型有提升作用,但在其他类型上反而降低性能。这些发现为依据文档结构与推理需求选择处理策略提供了实证指导。

原文摘要 · Abstract (English)

Document intelligence requires accurate text extraction and reliable reasoning over document content. We introduce \textbf{DISCO}, a \emph{Document Intelligence Suite for COmparative Evaluation}, that evaluates optical character recognition (OCR) pipelines and vision-language models (VLMs) separately on parsing and question answering across diverse document types, including handwritten text, multilingual scripts, medical forms, infographics, and multi-page documents. Our evaluation shows that performance varies substantially across tasks and document characteristics, underscoring the need for complexity-aware approach selection. OCR pipelines are generally more reliable for handwriting and for long or multi-page documents, where explicit text grounding supports text-heavy reasoning, while VLMs perform better on multilingual text and visually rich layouts. Task-aware prompting yields mixed effects, improving performance on some document types while degrading it on others. These findings provide empirical guidance for selecting document processing strategies based on document structure and reasoning demands.

文档理解OCR视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。