arXiv:2512.14050cs.CV2025-12

用多模态方法检测真实场景文字数据中的标签错误,提升文字识别准确率。

SELECT: Detecting Label Errors in Real-world Scene Text Data

  • 结合图像文本编码器与字符级分词器,处理可变长度标签和字符错误。
  • 在真实场景数据上检测标签错误效果优于现有方法,显著提升STR准确率。
  • 首次支持可变长度标签的错误检测,适合数据清洗与高精度文字识别任务。

我们提出SELECT(Scene tExt Label Errors deteCTion),一种利用多模态训练检测真实场景文字数据标签错误的新方法。通过图像-文本编码器和字符级分词器,该方法有效解决序列标签长度不一、标签错位及字符级错误问题,在准确率和实际应用价值上均优于现有方法。此外,我们引入基于相似性的序列标签扰动(SSLC),在训练中人为注入错误以模拟真实世界中的标注失误,不仅改变序列长度,还考虑字符间的视觉相似性。SELECT是首个成功处理可变长度标签的真实场景文字数据错误检测方法。实验表明,该方法能有效识别标签错误,并在真实场景文本数据集上提升文字识别(STR)准确率,展现显著实用价值。

原文摘要 · Abstract (English)

We introduce SELECT (Scene tExt Label Errors deteCTion), a novel approach that leverages multi-modal training to detect label errors in real-world scene text datasets. Utilizing an image-text encoder and a character-level tokenizer, SELECT addresses the issues of variable-length sequence labels, label sequence misalignment, and character-level errors, outperforming existing methods in accuracy and practical utility. In addition, we introduce Similarity-based Sequence Label Corruption (SSLC), a process that intentionally introduces errors into the training labels to mimic real-world error scenarios during training. SSLC not only can cause a change in the sequence length but also takes into account the visual similarity between characters during corruption. Our method is the first to detect label errors in real-world scene text datasets successfully accounting for variable-length labels. Experimental results demonstrate the effectiveness of SELECT in detecting label errors and improving STR accuracy on real-world text datasets, showcasing its practical utility.

文字识别数据清洗多模态标签错误

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。