arXiv:2605.10916cs.CVcs.AI2026-05

用可信度引导的扩散生成,提升低分辨率孟加拉复合字符识别准确率

Confidence-Guided Diffusion Augmentation for Enhanced Bangla Compound Character Recognition

  • 通过条件扩散模型与分类器引导生成高质量手写字符样本
  • 在AIBangla数据集上达到89.2%准确率,超越原有基准
  • 适合低资源语言文字识别研究者参考

由于复杂的字符结构、类内差异大以及高质量标注数据稀缺,手写孟加拉复合字符识别仍具挑战性。现有系统难以跨书写风格泛化,尤其对包含复杂连笔和变音符号的复合字符表现不佳。本文提出一种可信度引导的扩散增强框架,用于低分辨率孟加拉复合字符识别。该框架结合类别条件扩散建模与分类器引导,生成高质量手写样本。为提升生成质量,我们在扩散模型的U-Net主干中引入了增强型挤压-激励残差块,并设计基于可信度的过滤机制:预训练分类器作为质量门控,仅保留高类别一致性合成样本。将筛选后的合成图像与原始数据融合后,用于重训ResNet50、DenseNet121、VGG16及Vision Transformer等多类分类架构。在AIBangla复合字符数据集上的实验表明,各类模型均实现稳定提升,最优模型达89.2%分类准确率,显著超越先前发表的基准结果。结果证明,质量感知的扩散增强能有效提升低资源语种手写字符识别性能。

原文摘要 · Abstract (English)

Recognition of handwritten Bangla compound characters remains a challenging problem due to complex character structures, large intra-class variation, and limited availability of high-quality annotated data. Existing Bangla handwritten character recognition systems often struggle to generalize across diverse writing styles, particularly for compound characters containing intricate ligatures and diacritical variations. In this work, we propose a confidence-guided diffusion augmentation framework for low-resolution Bangla compound character recognition. Our framework combines class-conditional diffusion modeling with classifier guidance to synthesize high-quality handwritten compound character samples. To further improve generation quality, we introduce Squeeze-and-Excitation enhanced residual blocks within the diffusion model's U-Net backbone. We additionally propose a confidence-based filtering mechanism where pre-trained classifiers act as quality gates to retain only highly class-consistent synthetic samples. The filtered synthetic images are fused with the original training data and used to retrain multiple classification architectures. Experiments conducted on the AIBangla compound character dataset demonstrate consistent performance improvements across ResNet50, DenseNet121, VGG16, and Vision Transformer architectures. Our best-performing model achieves 89.2\% classification accuracy, surpassing the previously published AIBangla benchmark by a substantial margin. The results demonstrate that quality-aware diffusion augmentation can effectively enhance handwritten character recognition performance in low-resource script domains.

字符识别扩散模型低资源语言数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。