arXiv:2608.25140cs.CVcs.CL2026-08被引 1

用参考文本精准对齐手写稿行,自动标注历史阿拉伯文文献。

RefLAM: A Reference-Grounded Line Annotation Pipeline for Historical Arabic Manuscripts

论文配图:RefLAM: A Reference-Grounded Line Annotation Pipeline for Historical Arabic Manuscripts
图 1 · 摘自论文原文
  • 结合深度分割与多模态大模型,实现行级文本对齐。
  • 75倍提速,每小时处理3000行,99%以上对齐置信度达100。
  • 适合历史文献数字化与手写识别研究者使用。

现有手写阿拉伯文识别(HTR)训练数据构建方法要么依赖人工标注难以扩展,要么依赖自动OCR对齐但无法保证多文字、双区域(正文+边注)布局的准确性。本文提出RefLAM:一种将手稿图像与清晰转录文本转换为验证过的行级真值的流水线,兼顾人工监督。该方法融合深度学习页面分割模型、多模态大语言模型(MLLM)结构化OCR及无变音符模糊对齐引擎,将每行OCR结果锚定于参考文本连续段落,并提供0~100的字符级置信度。置信度100等价于归一化字符串完全一致(可信规则),在发布语料库中经验证无反例。评审者可信赖100分结果,仅需关注不确定对齐,大幅降低校验负担。在7本全页验证书籍上,效率相较人工提升75倍(3000 vs. 40行/小时);对另7本书应用同一标准,一周内保留16,533条置信度100的正文行,避免逐行修正。基于此,我们发布AraMS-28k:包含14部历史阿拉伯手稿、3,043页、27,971条正文行与629条边注行,含边界框、布局标签和191个边注插入锚点(占30.4%)。同时在AraMS-28k上微调预训练模型(如HATFormer),报告了字符错误率(CER),验证其对下游HTR训练的实际价值。

原文摘要 · Abstract (English)

Existing approaches to building line-level Arabic handwritten-text-recognition (HTR) training data either rely on fully manual annotation, which does not scale, or on automatic OCR-to-reference alignment methods not yet extended to multi-script, two-zone (main-plus-margin) manuscript layouts with a provable correctness guarantee. We present RefLAM (Reference-grounded Line Annotation for Manuscripts), a pipeline converting manuscript page images and clean transcriptions into validated, line-level ground truth without sacrificing human oversight. RefLAM couples a deep-learning page-segmentation model with a multimodal large language model (MLLM) for structured OCR and a diacritic-agnostic fuzzy alignment engine that grounds each OCR line in a contiguous span of the reference text, with a character-level confidence score in $[0,100]$. A perfect score is provably equivalent to character-for-character identity of the normalised strings (the Confidence-100 rule), verified with no counterexample across the released corpus. A reviewer can thus trust a perfect score, confirming most lines at a glance rather than retyping them, so annotation becomes triaged, with attention concentrated on uncertain alignments. Across 7 fully page-validated books we measured a 75$\times$ throughput gain over manual annotation (3,000 vs. 40 lines/hr); applying the same guarantee to 7 further books, we retained 16,533 confidence-100 main-text lines within one week, excluding sub-100 lines rather than manually correcting them. Using RefLAM, we release AraMS-28k: 14 historical Arabic manuscript books, 3,043 pages, and 27,971 main-text and 629 margin-line annotations with bounding boxes, layout labels, and insertion anchors for 191 margin entries (30.4%). We also finetune Muharaf-pretrained baselines (including HATFormer) on AraMS-28k and report CER results confirming its practical utility for downstream HTR training.

手写识别历史文献文本对齐阿拉伯文

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。