arXiv:2603.02789cs.CLcs.AI2026-03Conference of the …被引 1

强模型无需OCR也能精准提取文档信息,简化流程仍高效。

OCR or Not? Rethinking Document Information Extraction in the MLLMs Era with Real-World Large-Scale Datasets

  • 用大模型自动诊断错误模式,系统分析性能瓶颈。
  • 仅输入图像的MML模型表现接近带OCR的方案。
  • 精心设计提示和示例能进一步提升抽取准确率。

多模态大语言模型(MLLMs)提升了自然语言处理潜力,但其在文档信息抽取中的实际效果尚不明确。本文基于真实世界大规模数据集,对多种开箱即用的MLLMs在商业文档信息抽取任务上进行了大规模基准测试。为分析失败原因,提出一种基于大模型的自动化分层错误分析框架,可系统化诊断错误模式。结果表明,对于强大模型,纯图像输入即可达到与传统OCR+MLLM方案相当的效果;同时,精心设计的模式、示例和指令可进一步提升性能。本研究为文档信息抽取的实践优化提供了实用指导与关键洞察。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) enhance the potential of natural language processing. However, their actual impact on document information extraction remains unclear. In particular, it is unclear whether an MLLM-only pipeline--while simpler--can truly match the performance of traditional OCR+MLLM setups. In this paper, we conduct a large-scale benchmarking study that evaluates various out-of-the-box MLLMs on business-document information extraction. To examine and explore failure modes, we propose an automated hierarchical error analysis framework that leverages large language models (LLMs) to diagnose error patterns systematically. Our findings suggest that OCR may not be necessary for powerful MLLMs, as image-only input can achieve comparable performance to OCR-enhanced approaches. Moreover, we demonstrate that carefully designed schema, exemplars, and instructions can further enhance MLLMs performance. We hope this work can offer practical guidance and valuable insight for advancing document information extraction.

文档抽取多模态大模型无OCR

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。