arXiv:2505.10055cs.CVcs.AI2025-05被引 1

构建首个帕施图语光学字符识别基准,评测大模型表现。

PsOCR: Benchmarking Large Multimodal Models for Optical Character Recognition in Low-resource Pashto Language

  • 用百万级合成图像构建帕施图语OCR数据集
  • 闭源模型中Gemini表现最优,开源中Qwen-7B最佳
  • 为低资源语言文字识别提供可复用的基准

本文评估了大型多模态模型(LMMs)在低资源帕施图语光学字符识别(OCR)任务中的表现。由于帕施图语书写连笔且缺乏结构化数据集,自然语言处理面临挑战。为此,我们构建了包含一百万张带词、行和文档级别边界框标注的合成帕施图语OCR数据集PsOCR,覆盖1000种独特字体、颜色、图像尺寸与版式。从中选取1万张图像组成基准测试子集,用于评估七款开源模型(DeepSeek Janus、InternVL、MiniCPM、Florence、Qwen 3B/7B)及四款闭源模型(GPT-4o、Gemini、Claude、Grok)。实验表明,Gemini在所有模型中表现最佳,而开源模型中Qwen-7B领先。该工作揭示了当前LMM在帕施图语OCR中的能力与局限,为帕施图语及其他类似脚本(如阿拉伯语、波斯语、乌尔都语)的研究奠定了基础。PsOCR数据集已公开于https://github.com/zirak-ai/PashtoOCR。

原文摘要 · Abstract (English)

This paper evaluates the performance of Large Multimodal Models (LMMs) on Optical Character Recognition (OCR) in the low-resource Pashto language. Natural Language Processing (NLP) in Pashto faces several challenges due to the cursive nature of its script and a scarcity of structured datasets. To address this, we developed a synthetic Pashto OCR dataset, PsOCR, consisting of one million images annotated with bounding boxes at word, line, and document levels, suitable for training and evaluating models based on different architectures, including Convolutional Neural Networks (CNNs) and Transformers. PsOCR covers variations across 1,000 unique font families, colors, image sizes, and layouts. A benchmark subset of 10K images was selected to evaluate the performance of several LMMs, including seven open-source models: DeepSeek's Janus, InternVL, MiniCPM, Florence, and Qwen (3B and 7B), and four closed-source models: GPT-4o, Gemini, Claude, and Grok. Experimental results demonstrate that Gemini achieves the best performance among all models, whereas among open-source models, Qwen-7B stands out. This work provides an insightful assessment of the capabilities and limitations of current LMMs for OCR tasks in Pashto and establishes a foundation for further research not only in Pashto OCR but also for other similar scripts such as Arabic, Persian, and Urdu. PsOCR is available at https://github.com/zirak-ai/PashtoOCR.

OCR多模态低资源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。