arXiv:2502.09026cs.CV2025-02

用测试时自适应和编码规则提升钢坯号识别准确率

Billet Number Recognition Based on Test-Time Adaptation

  • 测试时通过降低模型熵实现无监督自适应
  • 结合编码规则过滤无效识别结果,提升准确率
  • 适合工业场景中复杂光照与破损文本识别

在钢铁钢坯生产过程中,实时识别运动中的钢坯号(机器打印或手写)至关重要。针对现有场景文字识别方法因图像畸变及训练与测试数据分布差异导致识别率低的问题,本文提出一种融合测试时自适应与先验知识的钢坯号识别方法。首先,将测试时自适应引入基于DB文本检测与SVTR文本识别的模型,在测试阶段通过最小化模型熵实现对测试数据分布的自适应,无需监督微调。其次,利用钢坯号编码规则作为先验知识,评估并替换不符合规则的识别结果。最后,改进CTC算法,引入先验知识验证机制以应对损坏字符识别瓶颈。在包含机器打印与手写钢坯号的真实数据集上的实验表明,各项评估指标显著提升,验证了方法的有效性。

原文摘要 · Abstract (English)

During the steel billet production process, it is essential to recognize machine-printed or manually written billet numbers on moving billets in real-time. To address the issue of low recognition accuracy for existing scene text recognition methods, caused by factors such as image distortions and distribution differences between training and test data, we propose a billet number recognition method that integrates test-time adaptation with prior knowledge. First, we introduce a test-time adaptation method into a model that uses the DB network for text detection and the SVTR network for text recognition. By minimizing the model's entropy during the testing phase, the model can adapt to the distribution of test data without the need for supervised fine-tuning. Second, we leverage the billet number encoding rules as prior knowledge to assess the validity of each recognition result. Invalid results, which do not comply with the encoding rules, are replaced. Finally, we introduce a validation mechanism into the CTC algorithm using prior knowledge to address its limitations in recognizing damaged characters. Experimental results on real datasets, including both machine-printed billet numbers and handwritten billet numbers, show significant improvements in evaluation metrics, validating the effectiveness of the proposed method.

文本识别测试时自适应工业视觉先验知识

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。