arXiv:2509.18081cs.CV2025-09EMNLP被引 2

用图谱词元化提升效率,实现高精度孟加拉手写文字识别。

GraDeT-HTR: A Resource-Efficient Bengali Handwritten Text Recognition System utilizing Grapheme-based Tokenizer and Decoder-only Transformer

  • 采用基于音素的分词器与解码器仅架构结合
  • 在多个基准数据集上达当前最优性能
  • 适合资源受限场景下的手写文本识别应用

尽管孟加拉语是全球第六大使用语言,其手写文字识别(HTR)系统仍严重不足。孟加拉文书写复杂,包含复合字母、变音符号及高度多样的手写风格,加之标注数据稀缺,使该任务尤为困难。我们提出GraDeT-HTR,一种基于音素感知解码器仅变压器架构的资源高效孟加拉手写文字识别系统。通过引入基于音素的分词器,显著提升解码器仅模型在孟加拉文上的识别准确率。模型在大规模合成数据上预训练,并在真实人工标注样本上微调,于多个基准数据集上实现当前最优性能。

原文摘要 · Abstract (English)

Despite Bengali being the sixth most spoken language in the world, handwritten text recognition (HTR) systems for Bengali remain severely underdeveloped. The complexity of Bengali script--featuring conjuncts, diacritics, and highly variable handwriting styles--combined with a scarcity of annotated datasets makes this task particularly challenging. We present GraDeT-HTR, a resource-efficient Bengali handwritten text recognition system based on a Grapheme-aware Decoder-only Transformer architecture. To address the unique challenges of Bengali script, we augment the performance of a decoder-only transformer by integrating a grapheme-based tokenizer and demonstrate that it significantly improves recognition accuracy compared to conventional subword tokenizers. Our model is pretrained on large-scale synthetic data and fine-tuned on real human-annotated samples, achieving state-of-the-art performance on multiple benchmark datasets.

手写识别中文处理轻量模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。