arXiv:2502.06172cs.CV2025-02被引 8

提出端到端手写印地语文字识别系统,解决多语言数据缺失与检测识别脱节问题。

PLATTER: A Page-Level Handwritten Text Recognition System for Indic Scripts

  • 将页面级手写文字识别分为文字检测与识别两阶段,独立优化各环节
  • 在10种印地语系语言上对比6个HTR模型,实现可复现的公平评估
  • 发布首个标注完善的印地语手写文本数据集CHIPS,支持检测与识别

近年来手写文字识别(HTR)领域涌现多种新模型,但因测试集不一致,难以进行公平比较。现有方法多聚焦字符或单词级识别,忽视了构建端到端手写文字识别流程所需的文字检测阶段,尤其对印地语系语言支持不足,主要受限于缺乏标注数据。本文提出一种页面级手写文字识别框架PLATTER,将其视为两阶段任务:先进行词级手写文字检测(HTD),再执行识别。该设计使各阶段问题可独立分析与优化。我们利用PLATTER评估无语言依赖的HTD模型,并在10种不同印地语系语言上对6个训练好的HTR模型进行一致性能对比,推动标准化评估。此外,我们发布了精心构建的印地语手写文本语料库CHIPS,包含页面级标注,覆盖检测与识别双重任务。同时开源代码与训练模型,以促进该方向的进一步研究。

原文摘要 · Abstract (English)

In recent years, the field of Handwritten Text Recognition (HTR) has seen the emergence of various new models, each claiming to perform competitively better than the other in specific scenarios. However, making a fair comparison of these models is challenging due to inconsistent choices and diversity in test sets. Furthermore, recent advancements in HTR often fail to account for the diverse languages, especially Indic languages, likely due to the scarcity of relevant labeled datasets. Moreover, much of the previous work has focused primarily on character-level or word-level recognition, overlooking the crucial stage of Handwritten Text Detection (HTD) necessary for building a page-level end-to-end handwritten OCR pipeline. Through our paper, we address these gaps by making three pivotal contributions. Firstly, we present an end-to-end framework for Page-Level hAndwriTTen TExt Recognition (PLATTER) by treating it as a two-stage problem involving word-level HTD followed by HTR. This approach enables us to identify, assess, and address challenges in each stage independently. Secondly, we demonstrate the usage of PLATTER to measure the performance of our language-agnostic HTD model and present a consistent comparison of six trained HTR models on ten diverse Indic languages thereby encouraging consistent comparisons. Finally, we also release a Corpus of Handwritten Indic Scripts (CHIPS), a meticulously curated, page-level Indic handwritten OCR dataset labeled for both detection and recognition purposes. Additionally, we release our code and trained models, to encourage further contributions in this direction.

手写识别印地语数据集端到端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。