arXiv:2605.26712cs.CV2026-05被引 1

构建多语言动态文本识别基准,助力真实场景模型评估与选型。

METATR: A Multilingual, Evolving Benchmark for Automatic Text Recognition

论文配图:METATR: A Multilingual, Evolving Benchmark for Automatic Text Recognition
图 1 · 摘自论文原文
  • 设计涵盖29种语言、多种文字和版式的多源文档数据集
  • 发现闭源模型性能最稳,但不同文字/布局间差异仍显著
  • 提供标准化评测流程,适合实际应用中模型对比与选型

真实世界文档的多样性和复杂性对自动文本识别(ATR)系统评估至关重要,尤其针对视觉大语言模型(vLLMs)。现有模型多在以英文为主的现代印刷体数据集上评估,难以反映实际应用场景。为此,本文提出METATR(v1.0),一个支持多语言、可演化的基准,用于评估ATR模型在多样化文档上的表现。该基准整合多个公开数据集,覆盖29种语言、多种书写系统与版式结构,并定义了统一的提示与归一化方法,建立动态可扩展的评估框架,确保结果可复现。我们测试了多种前沿模型(开源与闭源)。结果显示,尽管闭源模型整体表现更一致,但在不同文字和布局下仍存在显著性能波动。METATR为多语言ATR在真实条件下的评估提供了多维度、面向实践的工具,并可随领域发展持续更新。

原文摘要 · Abstract (English)

Benchmarks that reflect the diversity and complexity of real-world documents are essential for accurately evaluating Automatic Text Recognition (ATR) systems, especially Vision-Large Language Models (vLLMs). Although recent models demonstrate impressive performance, they are often evaluated on datasets containing modern, printed texts mostly written in English, which limits their relevance to many practical applications. Therefore, selecting a model for a specific use case requires evaluating it on data that matches the target documents. This highlights the importance of representative benchmarks for real-world applications. In this paper, we introduce METATR (v1.0), a multilingual, evolving benchmark designed to evaluate ATR models across a wide range of documents, facilitating meaningful model comparison and selection. The benchmark was designed to maximize diversity by including documents from various public collections. These documents cover 29 languages and include texts with multiple scripts and layouts. Beyond the dataset itself, METATR defines a standardized prompting and normalization methodology and establishes a dynamic evaluation framework. This approach is intended to produce reproducible results while remaining extensible over time. We evaluated a wide range of state-of-the-art systems, including open-source models and closed-source models. Results are reported across various dimensions, including performance at the dataset and language levels, robustness to handwritten documents, and computational efficiency. Our findings show that, although proprietary models achieve the most consistent performance, substantial variability persists across scripts and layouts. Overall, METATR provides a multidimensional, practitioner-oriented framework for assessing multilingual ATR in real-world conditions and tracking progress as the field evolves.

文本识别多语言评估基准vLLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。