arXiv:2502.14949cs.CVcs.AI2025-02ACL被引 22

构建首个覆盖9大领域36子类的阿拉伯文文档理解基准,推动阿拉伯文OCR技术发展。

KITAB-Bench: A Comprehensive Multi-Domain Benchmark for Arabic OCR and Document Understanding

  • 构建涵盖8809份样本的多领域阿拉伯文文档数据集
  • 视觉语言模型相比传统OCR平均降低60%字符错误率
  • 揭示复杂字体与表格识别短板,适合研究阿拉伯文文档分析者

随着检索增强生成(RAG)在文档处理中的广泛应用,鲁棒的文本识别对知识提取愈发关键。尽管英语等语言的OCR得益于大规模数据集和成熟基准,阿拉伯文OCR仍面临连笔书写、从右到左排版及复杂字体与书法特征等挑战。本文提出KITAB-Bench,一个全面的阿拉伯文OCR评估基准,包含9个主要领域、36个子领域,共8809个样本,覆盖手写文本、结构化表格及21种商业智能图表类型。结果显示,GPT-4o、Gemini、Qwen等视觉语言模型在字符错误率(CER)上比EasyOCR、PaddleOCR、Surya等传统方法平均提升60%。然而,当前模型在PDF转Markdown任务中表现仍弱,最佳模型Gemini-2.0-Flash仅达65%准确率,暴露出复杂字体、数字识别错误、单词拉伸和表格结构检测等难题。该工作建立了严谨的评估框架,有望推动阿拉伯文文档分析技术进步,缩小与英文OCR的性能差距。

原文摘要 · Abstract (English)

With the growing adoption of Retrieval-Augmented Generation (RAG) in document processing, robust text recognition has become increasingly critical for knowledge extraction. While OCR (Optical Character Recognition) for English and other languages benefits from large datasets and well-established benchmarks, Arabic OCR faces unique challenges due to its cursive script, right-to-left text flow, and complex typographic and calligraphic features. We present KITAB-Bench, a comprehensive Arabic OCR benchmark that fills the gaps in current evaluation systems. Our benchmark comprises 8,809 samples across 9 major domains and 36 sub-domains, encompassing diverse document types including handwritten text, structured tables, and specialized coverage of 21 chart types for business intelligence. Our findings show that modern vision-language models (such as GPT-4o, Gemini, and Qwen) outperform traditional OCR approaches (like EasyOCR, PaddleOCR, and Surya) by an average of 60% in Character Error Rate (CER). Furthermore, we highlight significant limitations of current Arabic OCR models, particularly in PDF-to-Markdown conversion, where the best model Gemini-2.0-Flash achieves only 65% accuracy. This underscores the challenges in accurately recognizing Arabic text, including issues with complex fonts, numeral recognition errors, word elongation, and table structure detection. This work establishes a rigorous evaluation framework that can drive improvements in Arabic document analysis methods and bridge the performance gap with English OCR technologies.

阿拉伯文OCR文档理解多领域基准视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。