融合CNN与Transformer提升阿拉伯手写体识别准确率
Bridging the Gap: Fusing CNNs and Transformers to Decode the Elegance of Handwritten Arabic Script
- 用CNN提取局部特征,Transformer捕捉长程依赖,协同处理变体
- 在IFN/ENIT数据集上实现96.38%字母识别率与97.22%位置识别率
- 适合需要高精度手写文字识别的文档数字化场景
阿拉伯手写体识别因字母形态动态变化和上下文依赖而极具挑战。本文提出一种结合卷积神经网络(CNN)与基于Transformer的架构的混合方法,评估了自定义及微调模型(包括EfficientNet-B7和Vision Transformer ViT-B16),并引入基于置信度融合的集成模型以整合两者优势。在IFN/ENIT数据集上,该集成模型实现96.38%的字母分类准确率与97.22%的位置分类准确率。结果表明CNN与Transformer具有互补性,展示了其联合应用在鲁棒阿拉伯手写识别中的潜力。本研究推动了OCR系统发展,为实际应用场景提供了可扩展的解决方案。
原文摘要 · Abstract (English)
Handwritten Arabic script recognition is a challenging task due to the script's dynamic letter forms and contextual variations. This paper proposes a hybrid approach combining convolutional neural networks (CNNs) and Transformer-based architectures to address these complexities. We evaluated custom and fine-tuned models, including EfficientNet-B7 and Vision Transformer (ViT-B16), and introduced an ensemble model that leverages confidence-based fusion to integrate their strengths. Our ensemble achieves remarkable performance on the IFN/ENIT dataset, with 96.38% accuracy for letter classification and 97.22% for positional classification. The results highlight the complementary nature of CNNs and Transformers, demonstrating their combined potential for robust Arabic handwriting recognition. This work advances OCR systems, offering a scalable solution for real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。