arXiv:2608.22366cs.CV2026-08

研究VLM在阿拉伯手稿OCR中的作用,发现条件校正能提升识别效果。

When Do VLMs Help Arabic Manuscript OCR? A Cross-Dataset Study

论文配图:When Do VLMs Help Arabic Manuscript OCR? A Cross-Dataset Study
图 1 · 摘自论文原文
  • 用OCR输出作为先验,引导VLM进行图像文本修正
  • 在历史手稿页级识别中,条件校正优于单独OCR或VLM
  • 对不同书写风格适应性不同,需根据文字可恢复性选择策略

视觉语言模型(VLMs)在文档理解中应用日益广泛,但在阿拉伯语及伊斯兰手稿识别中的作用仍缺乏研究。本文在涵盖历史手稿、老旧印刷书、清晰印刷体、多领域文档和手写的八组阿拉伯文本数据集上,评估了传统OCR、通用VLM、阿拉伯专用VLM以及基于OCR条件的VLM校正方法。结果表明,无单一方法在所有场景占优。在行级历史手稿中,VLM表现接近Tesseract;在页级手稿图像中表现更优;在多个设置中,基于OCR条件的校正器优于独立的OCR和VLM。核心发现是‘OCR先验可恢复性原则’:当OCR输出在视觉和文本上仍可恢复时,可为VLM提供锚点以优化图像内容。该方法在老旧印刷体、清晰印刷体、混合领域阿拉伯文和部分Naskh手稿中有效,但在马格里布体手稿和逼真学生手写中因字体不匹配或系统误导而性能下降。额外诊断显示阿拉伯VLM-OCR对变音符号、预处理、生成预算和重复循环敏感。研究支持一种自适应的OCR-VLM工作流,依据书写类型、先验可恢复性、长度诊断和故障指标分流处理页面。

原文摘要 · Abstract (English)

Vision-language models (VLMs) are increasingly being used for document understanding, yet their role in Arabic and Islamic manuscript recognition remains underexplored. To address such a gap in this paper, we evaluate traditional OCR, general-purpose VLMs, Arabic-specialized VLMs, and OCR-conditioned VLM correction across eight Arabic text datasets spanning historical manuscripts, aged printed books, clean print, multi-domain documents, and handwriting. The results show that no single approach dominates across setups. On line-level historical manuscripts, VLMs are close to Tesseract; on page-level manuscript images, they perform better; and in several settings, an OCR-conditioned corrector improves over both standalone OCR and standalone VLMs. The central finding is an OCR-prior recoverability principle: OCR conditioning helps when the OCR output remains visually and textually recoverable, providing anchors that the VLM can refine against the image. It improves recognition on aged print, clean print, mixed-domain Arabic, and some Naskh manuscripts, but degrades performance when the prior is script-mismatched or systematically misleading, as in Maghribi manuscripts and realistic student handwriting. Additional diagnostics show that Arabic VLM-OCR is sensitive to diacritics, preprocessing, generation budget, and repetition loops. These findings support an adaptive OCR-VLM workflow that routes pages according to script, OCR-prior recoverability, length diagnostics, and failure-mode indicators.

手稿识别VLMOCR阿拉伯语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。