arXiv:2504.05770cs.CVcs.AI2025-04

轻量级模型提升韩文字符识别准确率与速度

A Lightweight Multi-Module Fusion Approach for Korean Character Recognition

  • 融合笔画注意力与动态上下文编码,增强特征提取
  • 在多个基准上达到顶尖准确率,推理速度快30%以上
  • 适合实时和边缘设备部署,尤其适用于移动端

光学字符识别(OCR)在文档处理、车牌识别和智能监控中至关重要。然而,现有模型在真实场景中常因文本布局不规则、图像质量差、字符变化大及计算成本高而表现不佳。本文提出SDA-Net(笔画敏感注意力与动态上下文编码网络),一种轻量高效架构,专为鲁棒的单字符识别设计。该模型包含:(1) 双注意力机制,提升笔画级与空间特征提取能力;(2) 动态上下文编码模块,通过可学习门控机制自适应优化语义信息;(3) 类U-Net的特征融合策略,有效结合低层与高层特征;(4) 高度优化的轻量化主干网络,显著降低内存与计算开销。实验表明,SDA-Net在多个挑战性OCR基准上达到当前最优准确率,且推理速度大幅提升,非常适合实时与边缘端OCR系统部署。

原文摘要 · Abstract (English)

Optical Character Recognition (OCR) is essential in applications such as document processing, license plate recognition, and intelligent surveillance. However, existing OCR models often underperform in real-world scenarios due to irregular text layouts, poor image quality, character variability, and high computational costs. This paper introduces SDA-Net (Stroke-Sensitive Attention and Dynamic Context Encoding Network), a lightweight and efficient architecture designed for robust single-character recognition. SDA-Net incorporates: (1) a Dual Attention Mechanism to enhance stroke-level and spatial feature extraction; (2) a Dynamic Context Encoding module that adaptively refines semantic information using a learnable gating mechanism; (3) a U-Net-inspired Feature Fusion Strategy for combining low-level and high-level features; and (4) a highly optimized lightweight backbone that reduces memory and computational demands. Experimental results show that SDA-Net achieves state-of-the-art accuracy on challenging OCR benchmarks, with significantly faster inference, making it well-suited for deployment in real-time and edge-based OCR systems.

字符识别轻量模型OCR边缘计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。