arXiv:2608.19385cs.CV2026-08综述

让古阿拉伯手稿转录从识别升级为可审计的证据管理。

Beyond Recognition: Compact Multi-Domain Arabic Manuscript HTR with Candidate-Selection Analysis and Evidence-Preserving Review

  • 用多域自适应模型+证据保留流程,兼顾不同手写风格与不确定性。
  • 字符错误率降低25.3%,在多个数据集上实现领先性能。
  • 适合需要可追溯、可审查转录结果的研究者使用。

古阿拉伯手稿转录不仅是识别问题。一个可用的学术系统需应对手写体和版式变化,保留不确定读法,区分视觉证据与语言合理性,并记录研究者的最终判断。我们提出Phoenix——一个499万参数的CNN-BiLSTM-CTC识别器,以及围绕它的证据感知审核流程Athar。Phoenix通过文档感知回放、扩展至81符号的编码表及遗忘防护机制,在三个领域(档案、马格里布、历史手稿)间实现跨域适配,防止新领域提升以牺牲旧领域为代价。在预设保留集上的对比测试中,对10,594条Agapet文本,字符错误率(CER)从22.12%降至17.86%;对11,684条Omar文本,从17.72%降至11.84%;对164条TariMa文本,仅微升至10.72%。在两个大规模保留集上,加权字符错误率从19.98%降至14.93%,相对减少25.3%。开发诊断显示,Phoenix在四个可比领域中均达到最低错误率(未加权宏平均CER为9.59%)。N-best诊断揭示,束搜索与最优参考之间存在2.15点差距,而神经重排序、共识最小贝叶斯风险(MBR)、CTC后验质量估计及局部隐状态质量估计均未能恢复超过4%的差距。因此,Athar保留原始视觉读法,展示有限备选方案,保守使用局部语言模型,检索具有唯一性、模糊性或弃权状态的源文平行语料,并导出可审计的TEI与PAGE-XML格式记录。结果表明,手稿转录应被视为可审计的证据管理,而非无声的文本替换。

原文摘要 · Abstract (English)

Historical Arabic manuscript transcription is not only a recognition problem. A usable scholarly system must cope with shifting hands and layouts, preserve uncertain readings, distinguish visual evidence from linguistic plausibility, and record the researcher's final decision. We present Phoenix, a 4.99-million-parameter CNN-BiLSTM-CTC recognizer, and Athar, an evidence-aware review workflow built around it. Phoenix is adapted across archival, Maghrebi, and historical manuscript domains using document-aware replay, an expanded 81-symbol codec, and forgetting guards that reject checkpoints that improve a new domain at unacceptable cost to previous domains. In a pre-specified held-out comparison against the preceding checkpoint, frozen before evaluation and scored with greedy decoding and raw references, Phoenix reduced CER from 22.12% to 17.86% on 10,594 Agapet lines and from 17.72% to 11.84% on 11,684 Omar lines, while regressing from 10.39% to 10.72% on 164 TariMa lines. Across the two large held-out sets, character-weighted CER fell from 19.98% to 14.93%, a 25.3% relative error reduction. A separate same-protocol development diagnostic found the lowest CER for Phoenix on four of four comparable domains (9.59% unweighted macro CER). An N-best diagnostic revealed a 2.15-point oracle gap between beam decoding and Oracle@25, while neural text rerankers, consensus MBR, CTC-posterior quality estimation, and local pre-CTC hidden-state quality estimation recovered less than 4% of this gap. Athar therefore preserves the visual reading, exposes bounded alternatives, uses local language models conservatively, retrieves source parallels with unique, ambiguous, or abstain states, and exports auditable TEI and PAGE-XML records. The results support evaluating manuscript HTR as auditable evidence management rather than silent text replacement.

手稿转录证据管理多域适应可审计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。