提出人机协作框架,自动检测阿拉伯文字手写识别数据集中的标签错误。
A Human-in-the-Loop Label Error Detection Framework Applied to Arabic-Script HTR Datasets
- 基于卷积循环网络的字符错误率检测,先自动筛查可疑样本。
- 在多个数据集上识别出转录、分词等错误,最高准确率达90%。
- 适合需要高质量标注的阿拉伯文字手写识别研究者使用。
尽管近期取得进展,阿拉伯文字手写文本识别(HTR)仍落后于拉丁文字系统。部分原因在于数据集质量不高。为此,我们提出一种两阶段标签错误检测框架(CER-HV)。第一阶段(CER)基于卷积循环神经网络(CRNN)构建字符错误率噪声检测器;第二阶段(HV)由人工对第一阶段筛选出的可疑样本进行验证。该框架在多个阿拉伯文字数据集中成功识别出转录、分割、方向及非文本内容等标签错误,显著影响HTR性能。第一阶段检测准确率最高达90%(前50名样本)。我们的CRNN模型在六项评估数据集中的五项达到当前最优表现,其中在KHATT(阿拉伯语)上实现8.46%字符错误率(CER),PHTI(普什图语)为8.22%,Ajami为10.59%,Muharaf(阿拉伯语)为10.11%,均未经过数据清洗。在PHTD(波斯语)上建立11.3% CER新基准。应用CER-HV清理数据并重新训练后,评估CER最高提升1.8个百分点。尽管实验聚焦阿拉伯文字文档,该框架具备通用性,可推广至其他文本识别数据集。
原文摘要 · Abstract (English)
Despite recent advances, Handwritten Text Recognition (HTR) for Arabic-script languages still lags behind Latin-script HTR. Part of the problem is dataset quality. To help closing this gap, we propose a two-stage framework (CER-HV) for detecting label errors. Stage 1 (CER) is a Character-Error-Rate-based noise detector built on a Convolutional Recurrent Neural Network (CRNN) architecture. Stage 2 (HV) is the Human-In-The-Loop (HITL) Verification of noisy samples detected by the first stage. Applying the CER-HV framework on multiple Arabic-script datasets can identify samples with label errors including transcription, segmentation, orientation, and non-text content errors that can markedly affect HTR performance. These errors were identified by the first stage of the framework with up to 90percent (top-50) precision. We also show that our CRNN achieves state-of-the-art performance across five of the six evaluated datasets, reaching 8.46 percent Character Error Rate (CER) on KHATT (Arabic), 8.22 percent on PHTI (Pashto), 10.59 percent on Ajami, and 10.11% on Muharaf (Arabic), all without any data cleaning. We establish a new baseline of 11.3 percent CER on the PHTD (Persian) dataset. Applying CER-HV improves evaluation CER by up to 1.8 percentage points after dataset cleaning and retraining. Although our experiments focus on documents written in an Arabic-script language, the framework is general and can be applied to other text recognition datasets
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。