构建首个大规模多作者连续手写印地语数据集,用于提升低资源脚本识别性能。
DohaScript: A Large-Scale Multi-Writer Dataset for Continuous Handwritten Hindi Text
- 531位作者共同书写相同六首传统诗句,控制内容一致以分离书写风格差异
- 数据集含高质量图像与布局难度标注,支持跨作者泛化与风格分析任务
- 适合手写识别、作者识别及生成模型研究,尤其关注低资源脚本场景
尽管印地语有数亿使用者,但手写天城文文本在公开基准数据集中仍严重缺失。现有资源规模小,多聚焦孤立字符或短词,缺乏受控词汇内容与作者多样性,难以捕捉天城文手写中连笔、融合及结构复杂的特点——字符通过共享的shirorekha(横线)连接,并形成丰富连字。我们提出DohaScript,一个由531位独立贡献者收集的大规模多作者手写印地语文本数据集。该数据集设计为平行风格语料库,所有作者书写相同的六首传统印地语多哈(对句)。这种受控设计使可系统分析作者特异性差异,而与语言内容无关,支持手写识别、作者识别、风格分析与生成建模等任务。数据集附带非标识性人口统计元数据、基于客观清晰度与分辨率标准的质量筛选,以及页面级布局难度标注,便于分层基准测试。基线实验显示明显的质量区分与对未见作者的良好泛化能力,凸显数据集的可靠性与实用价值。DohaScript旨在成为低资源脚本环境下连续手写天城文研究的标准且可复现的基准。
原文摘要 · Abstract (English)
Despite having hundreds of millions of speakers, handwritten Devanagari text remains severely underrepresented in publicly available benchmark datasets. Existing resources are limited in scale, focus primarily on isolated characters or short words, and lack controlled lexical content and writer level diversity, which restricts their utility for modern data driven handwriting analysis. As a result, they fail to capture the continuous, fused, and structurally complex nature of Devanagari handwriting, where characters are connected through a shared shirorekha (horizontal headline) and exhibit rich ligature formations. We introduce DohaScript, a large scale, multi writer dataset of handwritten Hindi text collected from 531 unique contributors. The dataset is designed as a parallel stylistic corpus, in which all writers transcribe the same fixed set of six traditional Hindi dohas (couplets). This controlled design enables systematic analysis of writer specific variation independent of linguistic content, and supports tasks such as handwriting recognition, writer identification, style analysis, and generative modeling. The dataset is accompanied by non identifiable demographic metadata, rigorous quality curation based on objective sharpness and resolution criteria, and page level layout difficulty annotations that facilitate stratified benchmarking. Baseline experiments demonstrate clear quality separation and strong generalization to unseen writers, highlighting the dataset's reliability and practical value. DohaScript is intended to serve as a standardized and reproducible benchmark for advancing research on continuous handwritten Devanagari text in low resource script settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。