arXiv:2601.16713cs.CV2026-01被引 4

提出人机协作框架,自动检测阿拉伯文字手写识别数据集中的标签错误。

A Human-in-the-Loop Label Error Detection Framework Applied to Arabic-Script HTR Datasets

  • 基于卷积循环网络的字符错误率检测,先自动筛查可疑样本。
  • 在多个数据集上识别出转录、分词等错误,最高准确率达90%。
  • 适合需要高质量标注的阿拉伯文字手写识别研究者使用。

尽管近期取得进展,阿拉伯文字手写文本识别(HTR)仍落后于拉丁文字系统。部分原因在于数据集质量不高。为此,我们提出一种两阶段标签错误检测框架(CER-HV)。第一阶段(CER)基于卷积循环神经网络(CRNN)构建字符错误率噪声检测器;第二阶段(HV)由人工对第一阶段筛选出的可疑样本进行验证。该框架在多个阿拉伯文字数据集中成功识别出转录、分割、方向及非文本内容等标签错误,显著影响HTR性能。第一阶段检测准确率最高达90%(前50名样本)。我们的CRNN模型在六项评估数据集中的五项达到当前最优表现,其中在KHATT(阿拉伯语)上实现8.46%字符错误率(CER),PHTI(普什图语)为8.22%,Ajami为10.59%,Muharaf(阿拉伯语)为10.11%,均未经过数据清洗。在PHTD(波斯语)上建立11.3% CER新基准。应用CER-HV清理数据并重新训练后,评估CER最高提升1.8个百分点。尽管实验聚焦阿拉伯文字文档,该框架具备通用性,可推广至其他文本识别数据集。

原文摘要 · Abstract (English)

Despite recent advances, Handwritten Text Recognition (HTR) for Arabic-script languages still lags behind Latin-script HTR. Part of the problem is dataset quality. To help closing this gap, we propose a two-stage framework (CER-HV) for detecting label errors. Stage 1 (CER) is a Character-Error-Rate-based noise detector built on a Convolutional Recurrent Neural Network (CRNN) architecture. Stage 2 (HV) is the Human-In-The-Loop (HITL) Verification of noisy samples detected by the first stage. Applying the CER-HV framework on multiple Arabic-script datasets can identify samples with label errors including transcription, segmentation, orientation, and non-text content errors that can markedly affect HTR performance. These errors were identified by the first stage of the framework with up to 90percent (top-50) precision. We also show that our CRNN achieves state-of-the-art performance across five of the six evaluated datasets, reaching 8.46 percent Character Error Rate (CER) on KHATT (Arabic), 8.22 percent on PHTI (Pashto), 10.59 percent on Ajami, and 10.11% on Muharaf (Arabic), all without any data cleaning. We establish a new baseline of 11.3 percent CER on the PHTD (Persian) dataset. Applying CER-HV improves evaluation CER by up to 1.8 percentage points after dataset cleaning and retraining. Although our experiments focus on documents written in an Arabic-script language, the framework is general and can be applied to other text recognition datasets

手写识别数据清洗人机协作阿拉伯文

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。