构建首个覆盖149种语言的多语言文档解析基准,填补低资源语言评估空白。
MORE: A Multilingual Document Parsing Benchmark and Evaluation

- 覆盖149种语言,包含真实文档中的代码块、表格等复杂结构
- 基于模型辅助+人工校对的流程构建,数据真实可靠
- 适合研究多语言模型、跨语言信息提取的学者使用
多语言文档承载着丰富的区域文化、科学发现与历史记录。将其解析为机器可读的结构化格式,对挖掘全球知识至关重要。然而现有基准主要集中于英语、中文等高资源语言,难以验证模型在其他语言上的实际能力。尽管近期视觉-语言模型宣称支持数百种语言,但缺乏真实标注数据,无法实证其性能。为此,我们提出MORE——一个大规模多语言文档解析评估基准。MORE具有三大特点:(1) 超大规模:涵盖149种语言,是目前最语言多样化的基准;(2) 结构复杂性:不仅评估文本,还涵盖代码块、表格、目录等结构元素;(3) 数据真实性:所有样本均来自真实文档,通过模型辅助+人工精修流程生成。我们用MORE评估了前沿模型,建立了长尾语言的新性能基线,并验证了该基准在诊断模型真实场景表现方面的有效性。数据集将公开于https://github.com/zimoqingfeng/MORE。
原文摘要 · Abstract (English)
Multilingual documents encapsulate rich regional cultures, scientific discoveries, and historical records. Parsing this content into structured, machine-readable formats is critical for unlocking global knowledge. However, existing benchmarks predominantly focus on high-resource languages like English and Chinese, creating an evaluation blind spot concerning model performance on other languages. While recent Vision-Language Models (VLMs) claim support for hundreds of languages, the lack of ground truth makes it impossible to empirically verify these capabilities. To bridge this gap, we introduce MORE, a large-scale benchmark designed for multilingual document parsing evaluation. MORE distinguishes itself through three key dimensions: (1) Unprecedented Scale: It covers 149 languages, making it the most linguistically diverse benchmark to date; (2) Structural Complexity: Unlike previous works, it extends evaluation beyond plain text to include structural elements such as code blocks, tables, and catalogs; and (3) Data Authenticity: All samples are curated from real-world documents via a model-assisted, human-refined annotation pipeline. We evaluate state-of-the-art models using MORE, establishing new performance baselines for long-tail languages and validating the benchmark's effectiveness in diagnosing model capabilities in realistic, diverse scenarios. The MORE dataset will be available at https://github.com/zimoqingfeng/MORE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。