arXiv:2604.22515cs.CVcs.LG2026-04中稿 · publication in the…

提升历史阿拉伯手稿作者识别准确率,支持文化遗产研究

Different Strokes for Different Folks: Writer Identification for Historical Arabic Manuscripts

  • 用带注意力的CNN模型识别手稿作者,处理罕见双作者情况
  • 线级评估达99.05%准确率,页级分离评估仍达78.61%准确率
  • 首次提供线级与页级双评估基准,适合历史文献研究者

手写阿拉伯古籍承载了阿拉伯世界的思想与文化遗产,作者识别有助于溯源、真伪鉴定和历史分析。基于Muharaf古籍数据集,我们评估了从单行图像中进行作者识别的表现,并首次报告了线级与页不相交两种评估协议下的基准结果。原数据集仅28.00%的行有作者标签,我们手动验证并扩充至86.75%(21,249行/24,495行),修正不一致并剔除非手写内容,最终保留18,987行(77.51%)。提出一种结合注意力机制的CNN模型,用于闭集作者识别,包括将罕见双作者行建模为复合类别。对比十四种配置,分析不同特征提取器与训练策略。为评估对未见页面的泛化能力,采用页不相交协议——每页所有行归入同一划分。在线级协议下,微调DenseNet201+注意力模型达到99.05% Top-1准确率、99.73% Top-5准确率和97.44% F1分数;在更具挑战性的页不相交协议下,最优结果为78.61% Top-1准确率、87.79% Top-5准确率和66.55% F1分数,量化了页面线索的影响。通过扩展标注子集并报告双协议,为历史学家与语言学家提供了更清晰的基准与实用资源。代码与实现细节已在GitHub公开。

原文摘要 · Abstract (English)

Handwritten Arabic manuscripts preserve the Arab world's intellectual and cultural heritage, and writer identification supports provenance, authenticity verification, and historical analysis. Using the Muharaf dataset of historical Arabic manuscripts, we evaluate writer identification from individual line images and, to the best of our knowledge, provide the first baselines reported under both line-level and page-disjoint evaluation protocols. Since the dataset is only partially labeled for writer identification, we manually verified and expanded writer labels in the public portion from 6,858 (28.00%) to 21,249 lines (86.75%) out of 24,495 line images, correcting inconsistencies and removing non-handwritten text. After further filtering, we retained 18,987 lines (77.51%). We propose a Convolutional Neural Network (CNN)-based model with attention mechanisms for closed-set writer identification, including rare two-writer lines modeled as composite writer-pair classes. We benchmark fourteen configurations and conduct ablations across different feature extractors and training regimes. To assess generalization to unseen pages, the page-disjoint protocol assigns all lines from each page to a single split. Under the line-level protocol, a fine-tuned DenseNet201 with attention achieves 99.05% Top-1 accuracy, 99.73% Top-5 accuracy, and 97.44% F1-score. Under the more challenging page-disjoint protocol, the best observed results are 78.61% Top-1 accuracy, 87.79% Top-5 accuracy, and 66.55% F1-score, thus quantifying the impact of page-level cues. By expanding the Muharaf dataset's labeled subset and reporting both protocols, we provide a clearer benchmark and a practical resource for historians and linguists engaged with culturally and historically significant documents. The code and implementation details are available on GitHub.

手稿识别历史文献深度学习阿拉伯文

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。