arXiv:2603.28130cs.CVcs.AI2026-03被引 4

首个面向多语言真实文档的解析基准,揭示开源模型在复杂场景下性能严重下降。

MDPBench: A Benchmark for Multilingual Document Parsing in Real-World Scenarios

  • 构建跨17种语言、多种书写系统的真实文档图像数据集
  • 开源模型在拍摄文档上平均降17.8%,非拉丁文降14.0%
  • 适合关注多语言文档识别与公平性研究的开发者

我们提出多语言文档解析基准(MDPBench),这是首个针对多语言数字与照片文档解析的系统性评测基准。尽管文档解析技术已取得显著进展,但主要集中于少数主流语言的干净数字文档。目前尚无统一基准评估模型在多样书写系统和低资源语言中的表现。MDPBench包含3,400张涵盖17种语言、多种书写系统及不同拍摄条件的文档图像,采用专家模型标注、人工校对与人类验证的严格流程生成高质量标注。为确保公平比较并防止数据泄露,数据集划分为公开与私有测试集。对开源与闭源模型的全面评估发现:闭源模型(如Gemini3-Pro)表现相对稳健,而开源模型在非拉丁文和真实拍摄文档上出现显著性能下降,分别平均降低17.8%和14.0%。结果揭示了语言与场景间的显著性能差异,指明构建更具包容性与部署可行性的解析系统的方向。代码与数据见https://github.com/Yuliang-Liu/MultimodalOCR。

原文摘要 · Abstract (English)

We introduce Multilingual Document Parsing Benchmark, the first benchmark for multilingual digital and photographed document parsing. Document parsing has made remarkable strides, yet almost exclusively on clean, digital, well-formatted pages in a handful of dominant languages. No systematic benchmark exists to evaluate how models perform on digital and photographed documents across diverse scripts and low-resource languages. MDPBench comprises 3,400 document images spanning 17 languages, diverse scripts, and varied photographic conditions, with high-quality annotations produced through a rigorous pipeline of expert model labeling, manual correction, and human verification. To ensure fair comparison and prevent data leakage, we maintain separate public and private evaluation splits. Our comprehensive evaluation of both open-source and closed-source models uncovers a striking finding: while closed-source models (notably Gemini3-Pro) prove relatively robust, open-source alternatives suffer dramatic performance collapse, particularly on non-Latin scripts and real-world photographed documents, with an average drop of 17.8% on photographed documents and 14.0% on non-Latin scripts. These results reveal significant performance imbalances across languages and conditions, and point to concrete directions for building more inclusive, deployment-ready parsing systems. Source available at https://github.com/Yuliang-Liu/MultimodalOCR.

文档解析多语言真实场景基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。