首个开源摩洛哥方言手写体识别模型,用小模型实现高精度
AtlasOCR: Building the First Open-Source Darija OCR Model with Vision Language Models
- 用30亿参数视觉语言模型微调,结合合成与真实数据训练
- 在自建数据集上达到领先性能,优于更大模型
- 适合低资源语言、多模态研究者及阿拉伯语应用开发者
摩洛哥方言(Darija)虽富含视觉内容,却缺乏专用光学字符识别(OCR)工具。本文提出AtlasOCR,首个基于30亿参数视觉语言模型(VLM)微调的开源Darija OCR模型。我们构建了独特的Darija专属数据集,融合自研OCRSmith库生成的合成数据与精心采集的真实世界数据,并采用高效微调策略,使用QLoRA与Unsloth对Qwen2.5-VL 3B进行训练。通过全面的消融实验优化关键超参数。在新构建的AtlasOCRBench和现有KITAB-Bench上的评估显示,其性能达到业界领先水平,超越更大模型,展现出对Darija及标准阿拉伯语OCR任务的强鲁棒性与泛化能力。
原文摘要 · Abstract (English)
Darija, the Moroccan Arabic dialect, is rich in visual content yet lacks specialized Optical Character Recognition (OCR) tools. This paper introduces AtlasOCR, the first open-source Darija OCR model built by fine-tuning a 3B parameter Vision Language Model (VLM). We detail our comprehensive approach, from curating a unique Darija-specific dataset leveraging both synthetic generation with our OCRSmith library and carefully sourced real-world data, to implementing efficient fine-tuning strategies. We utilize QLoRA and Unsloth for parameter-efficient training of Qwen2.5-VL 3B and present comprehensive ablation studies optimizing key hyperparameters. Our evaluation on the newly curated AtlasOCRBench and the established KITAB-Bench demonstrates state-of-the-art performance, challenging larger models and highlighting AtlasOCR's robustness and generalization capabilities for both Darija and standard Arabic OCR tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。