arXiv:2608.26921cs.CVcs.CL2026-08被引 1

首个公开的阿拉伯古籍行级数据集,含边注与插入锚点标注。

AraMS-28k: The Largest Publicly Released Line-Level Dataset of Historical Arabic Manuscripts with Margin and Insertion-Anchor Annotations

论文配图:AraMS-28k: The Largest Publicly Released Line-Level Dataset of Historical Arabic Manuscripts with Margin and Insertion-Anchor Annotations
图 1 · 摘自论文原文
  • 构建基于参考文本对齐的标注流程,实现多模态大模型自动校验+人工审核。
  • 包含28,600行标注文本,其中629行为带明确插入锚点的边注。
  • 支持手写体、印刷体混合识别,适用于古籍文字识别与阅读顺序恢复研究。

我们推出AraMS-28k,目前最大的公开历史阿拉伯手稿行级数据集,涵盖14部书、3,043页和28,600个标注文本行(27,971行正文,629行边注)。其中13部为手抄本,覆盖纳斯赫、鲁卡阿和马格里布三种书写传统,1部为石印印刷版以增强格式多样性。每行标注为主文或边注,边注若在正文中存在明确对应位置,则进一步标注插入锚点,实现行级粒度的非线性阅读顺序恢复——据我们所知,这是首个针对历史阿拉伯手稿语料库发布的此类标注。因参考转录为全元音化而手稿通常无变音符号,我们为每行提供原始带变音转录及变音标准化版本。数据集通过RefLAM参考对齐标注流程构建,该流程将多模态大模型OCR与独立获取的清晰转录对齐,并经人工复核,结合自动验证与专家监督。我们详述构建与质控过程,介绍标注方案,报告语料库及各书级别的统计数据,并使用Kraken和HATFormer提供基线手写体识别结果,包括从分布内页面到完全未见书籍的跨书写风格泛化能力梯度。AraMS-28k以CC BY-NC-SA 4.0协议发布,包含页图像、行级标注及固定训练/验证/测试划分,支持阿拉伯古籍识别、版面分析与阅读顺序恢复的可复现研究。

原文摘要 · Abstract (English)

We introduce AraMS-28k, the largest publicly released line-level dataset of genuine historical Arabic manuscripts, comprising 14 books, 3,043 pages, and 28,600 annotated text lines (27,971 main-text, 629 margin). Thirteen books are hand-copied manuscripts spanning three script traditions -- Naskh, Ruq'ah, and Maghrebi -- and one is a lithographed printed edition included to broaden format diversity. Each line is labelled as main-text or margin, and margin lines that have an unambiguous attachment point in the main text are further annotated with an insertion anchor, recovering the manuscript's true non-linear reading order at line-level granularity -- to our knowledge the first such annotation released for a historical Arabic manuscript corpus. Because reference transcriptions are fully vocalised while manuscript hands are typically undiacritised, we release both the raw diacritised transcription and a diacritic-normalised counterpart for every line. The dataset was constructed with RefLAM, a reference-grounded annotation pipeline that aligns multimodal-LLM OCR against independently sourced clean transcriptions and routes every line through human review, combining automatic verification with expert oversight. We describe the construction and quality-control process, present the annotation schema, report dataset statistics at both the corpus and per-book level, and provide baseline HTR results using Kraken and HATFormer, including a cross-script generalisation gradient from in-distribution pages to fully unseen books. AraMS-28k is released with page images, line-level annotations, and fixed train/val/test splits under CC BY-NC-SA 4.0 to support reproducible research on Arabic manuscript recognition, layout analysis, and reading-order recovery.

古籍识别手写体标注数据集阅读顺序

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。