arXiv:2509.18174cs.CVcs.CL2025-09被引 4

专为阿拉伯文文档设计的视觉语言模型,显著提升文本识别准确率。

Baseer: A Vision-Language Model for Arabic Document-to-Markdown OCR

  • 基于合成与真实文档数据,微调预训练多模态大模型以适配阿拉伯文。
  • 在新基准上实现0.25的词错误率,超越现有开源与商用方案。
  • 适合需要高精度阿拉伯文文档识别的研究者与开发者使用。

阿拉伯文文档的OCR任务因连笔书写、字体多样、元音符号和从右到左的排版而极具挑战性。尽管现代多模态大模型在高资源语言上取得进展,但其在阿拉伯文上的表现仍有限。本文提出Baseer,一个专门针对阿拉伯文文档OCR优化的视觉语言模型。该模型利用大规模合成与真实文档数据集,采用仅解码器的微调策略,在保留通用视觉特征的同时适配预训练多模态大模型。我们还构建了Misraj-DocOCR,一个高质量、专家验证的基准测试集,用于严格评估阿拉伯文OCR系统。实验表明,Baseer显著优于现有开源与商业解决方案,在词错误率(WER)上达到0.25,确立了阿拉伯文文档OCR的新基准。结果表明,对通用多模态大模型进行领域特定适配具有显著优势,为形态丰富的语言如阿拉伯文提供了高精度OCR的可靠基线。

原文摘要 · Abstract (English)

Arabic document OCR remains a challenging task due to the language's cursive script, diverse fonts, diacritics, and right-to-left orientation. While modern Multimodal Large Language Models (MLLMs) have advanced document understanding for high-resource languages, their performance on Arabic remains limited. In this work, we introduce Baseer, a vision-language model fine-tuned specifically for Arabic document OCR. Leveraging a large-scale dataset combining synthetic and real-world documents, Baseer is trained using a decoder-only fine-tuning strategy to adapt a pre-trained MLLM while preserving general visual features. We also present Misraj-DocOCR, a high-quality, expert-verified benchmark designed for rigorous evaluation of Arabic OCR systems. Our experiments show that Baseer significantly outperforms existing open-source and commercial solutions, achieving a WER of 0.25 and establishing a new state-of-the-art in the domain of Arabic document OCR. Our results highlight the benefits of domain-specific adaptation of general-purpose MLLMs and establish a strong baseline for high-accuracy OCR on morphologically rich languages like Arabic.

文档OCR阿拉伯文多模态模型视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。