arXiv:2412.16119cs.LGcs.CV2024-12被引 7

评测大模型对低资源文字的识别能力,发现现有方法仍有明显短板。

Deciphering the Underserved: Benchmarking LLM OCR for Low-Resource Scripts

  • 用2520张带多种干扰的图像测试GPT-4o对乌尔都语等语言的识别
  • 零样本识别在复杂语言上准确率低,尤其在模糊或小字号时更差
  • 呼吁建立标注数据集,推动低资源语言的可访问性研究

本研究评估大型语言模型(如GPT-4o)在低资源脚本(如乌尔都语、阿尔巴尼亚语、塔吉克语)中的光学字符识别(OCR)潜力,以英语为基准。基于包含2,520张图像的精心构建数据集,模拟文本长度、字体大小、背景色和模糊度等现实挑战。结果表明,零样本的LLM OCR在语言结构复杂的脚本中表现受限,凸显了标注数据集与微调模型的必要性。研究强调了弥合文本数字化中语言可及性差距的紧迫性,为实现包容性强健的低资源语言OCR解决方案铺平道路。

原文摘要 · Abstract (English)

This study investigates the potential of Large Language Models (LLMs), particularly GPT-4o, for Optical Character Recognition (OCR) in low-resource scripts such as Urdu, Albanian, and Tajik, with English serving as a benchmark. Using a meticulously curated dataset of 2,520 images incorporating controlled variations in text length, font size, background color, and blur, the research simulates diverse real-world challenges. Results emphasize the limitations of zero-shot LLM-based OCR, particularly for linguistically complex scripts, highlighting the need for annotated datasets and fine-tuned models. This work underscores the urgency of addressing accessibility gaps in text digitization, paving the way for inclusive and robust OCR solutions for underserved languages.

OCR大模型低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。