arXiv:2606.19139cs.CVcs.CL2026-06

首个面向历史乌尔都手写体的基准数据集,助力古籍文字识别

Urdu Katib Handwritten Dataset: A Historical Document Dataset for Offline Urdu Handwritten Text Recognition with CRNN-Based Baseline Evaluation

论文配图:Urdu Katib Handwritten Dataset: A Historical Document Dataset for Offline Urdu Handwritten Text Recognition with CRNN-Based Baseline Evaluation
图 1 · 摘自论文原文
  • 构建首个基于历史文书的乌尔都手写文本数据集
  • CNN-BGRU-CTC模型表现最优,字符错误率低
  • 适合古籍数字化与手写识别研究者使用

自动手写文本识别(HTR)本身极具挑战性,尤其在连笔书写中更为复杂。尽管已有诸多关于连笔文字的研究,但针对乌尔都语手写文字识别(UHTR)的工作仍相对有限,主要因该文字体系的独特性以及缺乏基准数据集。为此,本文提出首个专门从历史时期卡提布(Katibs)书写的文献中收集的离线乌尔都手写文本行数据集——乌尔都卡提布手写数据集(UKHD)。该数据集涵盖多种平尖笔书写风格,体现纳斯塔里克书法特点。同时,评估了多种基于CRNN的混合模型在乌尔都卡提布手写识别(UKHR)中的性能。结果表明,CNN-BGRU-CTC模型表现更优,具备较低的字符错误率(CER)和词错误率(WER)。本研究旨在推动学术界发展鲁棒的识别系统,以保护乌尔都手写文献。

原文摘要 · Abstract (English)

Automatic Handwritten Text Recognition (HTR) is inherently a challenging task, and its complexity is further increased when dealing with cursive scripts. Although significant efforts have been made on various cursive scripts, research regarding Urdu Handwritten Text Recognition (UHTR) has been relatively limited. This lag of research is primarily due to the unique challenges posed by its script, and the scarcity and unavailability of benchmark datasets. Therefore, to advance research in UHTR, this study presents a specialized real dataset called the Urdu Katib Handwritten Dataset (UKHD). To the best of our knowledge, this is the first offline Urdu handwritten text lines dataset specifically curated from the materials written by Katibs in historical times. It encompasses a diverse range of flat nib writing variations in the Nastalique calligraphic style. Additionally, the effectiveness of different CRNN-based hybrid models has been evaluated to identify the optimal architecture for Urdu Katib Handwriting Recognition (UKHR). Among the analyzed models, the CNN-BGRU-CTC model showed more robust performance, with low Character Error Rate (CER) and Word Error Rate (WER). This research work aims to support and encourage the research community in developing a robust recognition system for preserving Urdu handwritten literature.

手写识别古籍数字化数据集乌尔都语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。