arXiv:2602.07015cs.CVeess.IV2026-02被引 1

用混合模型提升孟加拉纸币识别准确率,兼顾实时性与低资源消耗。

Robust and Real-Time Bangladeshi Currency Recognition: A Dual-Stream MobileNet and EfficientNet Approach

  • 融合MobileNetV3-Large与EfficientNetB0提取特征,兼顾速度与精度。
  • 在复杂背景中仍达92.84%准确率,综合数据集上达94.98%。
  • 适配手机端部署,支持视障人士独立辨识纸币,防欺诈。

精准的货币识别对辅助技术至关重要,尤其对依赖他人识别钞票的视障人士而言。为应对这一挑战,我们构建了一个包含受控环境与真实场景的孟加拉国纸币新数据集,确保更全面、多样化的覆盖。为进一步增强数据集鲁棒性,我们整合了四个额外数据集,包括公开基准,以涵盖多种复杂情况并提升模型泛化能力。针对现有识别模型的局限,提出一种新型混合卷积神经网络架构,结合MobileNetV3-Large与EfficientNetB0实现高效特征提取,并采用多层感知机(MLP)分类器,在保持低计算成本的同时提升性能,适用于资源受限设备。实验结果表明,该模型在受控数据集上达到97.95%准确率,在复杂背景中达92.84%,综合所有数据集时达94.98%。通过五折交叉验证和七项指标(准确率、精确率、召回率、F1-score、Cohen's Kappa、MCC、AUC)进行充分评估。同时引入LIME与SHAP等可解释性AI方法,提升系统透明度与可解释性。

原文摘要 · Abstract (English)

Accurate currency recognition is essential for assistive technologies, particularly for visually impaired individuals who rely on others to identify banknotes. This dependency puts them at risk of fraud and exploitation. To address these challenges, we first build a new Bangladeshi banknote dataset that includes both controlled and real-world scenarios, ensuring a more comprehensive and diverse representation. Next, to enhance the dataset's robustness, we incorporate four additional datasets, including public benchmarks, to cover various complexities and improve the model's generalization. To overcome the limitations of current recognition models, we propose a novel hybrid CNN architecture that combines MobileNetV3-Large and EfficientNetB0 for efficient feature extraction. This is followed by an effective multilayer perceptron (MLP) classifier to improve performance while keeping computational costs low, making the system suitable for resource-constrained devices. The experimental results show that the proposed model achieves 97.95% accuracy on controlled datasets, 92.84% on complex backgrounds, and 94.98% accuracy when combining all datasets. The model's performance is thoroughly evaluated using five-fold cross-validation and seven metrics: accuracy, precision, recall, F1-score, Cohen's Kappa, MCC, and AUC. Additionally, explainable AI methods like LIME and SHAP are incorporated to enhance transparency and interpretability.

图像识别移动设备可解释性视障辅助

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。