构建古埃及语手写文本识别数据集,助力濒危语言文字数字化
A Text Recognition Dataset from Sahidic Coptic Ancient Manuscripts

- 基于数字化古籍构建线级手写文本数据集
- 涵盖墨迹褪色、纸张劣化等真实历史文献问题
- 为低资源语言文字识别提供基准测试参考
本文聚焦低资源场景下的手写文本识别(HTR),针对代表性不足的语言、罕见书写系统及历史文献常见的视觉退化问题。我们提出了SCAM(Sahidic Coptic Ancient Manuscripts)数据集,该数据集源自用已消亡的萨希迪克科普特方言书写的古代手稿,包含跨图书馆采集的异质成像条件和典型的文献退化特征,如墨迹褪色、渗透穿透与材质老化。除视觉复杂性外,萨希迪克科普特语本身因资源稀缺、字母罕见及方言特有的变音符号而带来显著语言挑战。为推动低资源HTR研究,我们基于不同范式的前沿方法进行基准测试,揭示了当前模型在主流现代文本与真实历史低资源场景间的性能差距,为未来研究提供了重要参考。
原文摘要 · Abstract (English)
In this work, we target Handwritten Text Recognition (HTR) in low-resource scenarios, which arise from underrepresented languages, rare scripts, and degraded visual conditions typical of historical documents. We introduce SCAM (Sahidic Coptic Ancient Manuscripts), a new line-level dataset built from digitized ancient manuscripts written in the extinct Sahidic Coptic dialect. The dataset reflects a realistic and challenging setting, as it combines heterogeneous acquisition conditions across libraries with typical manuscript degradations such as ink fading, bleed-through, and material deterioration. In addition to visual complexity, SCAM poses significant linguistic challenges due to the scarcity of resources for Sahidic Coptic, its uncommon alphabet, and dialect-specific diacritics. To support research in low-resource HTR, we benchmark several state-of-the-art approaches based on different paradigms, highlighting their limitations and strengths in this setting. Our results underline the gap between current HTR performance on well-resourced modern scripts and historically grounded, low-resource scenarios, thus providing a reference point for future developments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。