arXiv:2509.05019cs.CV2025-09被引 5

用轻量模型和迁移学习提升阿拉伯手写字符识别准确率

Leveraging Transfer Learning and Mobile-enabled Convolutional Neural Networks for Improved Arabic Handwritten Character Recognition

  • 用迁移学习+轻量CNN模型,解决数据少算力高难题
  • 全微调策略下最高达99%准确率,MobileNet表现最优
  • 适合移动端部署,尤其适合资源有限的阿拉伯语识别场景

本研究探索将迁移学习(TL)与移动端轻量卷积神经网络(MbNets)结合,以提升阿拉伯手写字符识别(AHCR)性能。针对计算需求大、数据稀缺等问题,评估了三种迁移学习策略(全微调、部分微调、从头训练)在四种轻量级MbNets(MobileNet、SqueezeNet、MnasNet、ShuffleNet)上的表现。实验在三个基准数据集(AHCD、HIJJA、IFHCDB)上进行。MobileNet始终表现最佳,兼具高精度、强鲁棒性和高效性;ShuffleNet在全微调下泛化能力突出。IFHCDB数据集表现最优,全微调下使用MnasNet达到99%准确率,表明其适合稳健识别。AHCD数据集取得97%准确率(ShuffleNet),而HIJJA因变异性大,最高仅达92%(ShuffleNet)。全微调整体表现最佳,兼顾精度与收敛速度,部分微调表现较差。结果表明,结合迁移学习与轻量网络可实现高效低资源的阿拉伯手写识别,为后续优化与应用奠定基础。未来工作将聚焦架构改进、数据集特征分析、数据增强与敏感性分析。

原文摘要 · Abstract (English)

The study explores the integration of transfer learning (TL) with mobile-enabled convolutional neural networks (MbNets) to enhance Arabic Handwritten Character Recognition (AHCR). Addressing challenges like extensive computational requirements and dataset scarcity, this research evaluates three TL strategies--full fine-tuning, partial fine-tuning, and training from scratch--using four lightweight MbNets: MobileNet, SqueezeNet, MnasNet, and ShuffleNet. Experiments were conducted on three benchmark datasets: AHCD, HIJJA, and IFHCDB. MobileNet emerged as the top-performing model, consistently achieving superior accuracy, robustness, and efficiency, with ShuffleNet excelling in generalization, particularly under full fine-tuning. The IFHCDB dataset yielded the highest results, with 99% accuracy using MnasNet under full fine-tuning, highlighting its suitability for robust character recognition. The AHCD dataset achieved competitive accuracy (97%) with ShuffleNet, while HIJJA posed significant challenges due to its variability, achieving a peak accuracy of 92% with ShuffleNet. Notably, full fine-tuning demonstrated the best overall performance, balancing accuracy and convergence speed, while partial fine-tuning underperformed across metrics. These findings underscore the potential of combining TL and MbNets for resource-efficient AHCR, paving the way for further optimizations and broader applications. Future work will explore architectural modifications, in-depth dataset feature analysis, data augmentation, and advanced sensitivity analysis to enhance model robustness and generalizability.

手写识别迁移学习轻量模型阿拉伯语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。