arXiv:2608.26500cs.CVcs.LG2026-08综述被引 11

梳理十年来OCR模型演进,揭示多语言与复杂文本识别的突破与瓶颈

Systematic Literature Review of Machine Learning Models and Applications for Text Recognition

论文配图:Systematic Literature Review of Machine Learning Models and Applications for Text Recognition
图 1 · 摘自论文原文
  • 按PRISMA标准分析97篇论文,追踪AI模型在文本识别中的发展脉络
  • 发现自监督学习与多模态融合显著提升多语言和手写体识别效果
  • 适合关注OCR技术演进、跨语言处理及工业落地的研究者参考

基于系统性文献综述(PRISMA)指南,本文对2015年1月至2025年1月间发表的97篇研究进行深度分析,全面评估过去十年中机器学习模型在文本识别领域的进展。研究聚焦于模型架构演变、应用领域扩展、数据类型多样性、语言覆盖范围及现存挑战。结果表明,现代OCR已能有效处理结构化与非结构化文本、场景文字识别及多语言场景。然而,低资源语言数据匮乏、手写体差异大、字符视觉相似性高以及实时应用限制仍是未解难题。文中提出自监督学习、多模态AI、自动化机器学习(AutoML)、AI辅助后处理、微型机器学习(TinyML)及联合语料库构建等前瞻性策略,以提升识别精度并推动工业级实时应用。本研究为未来方向提供重要指引,奠定该领域研究基础。

原文摘要 · Abstract (English)

Optical Character Recognition (OCR) for text recognition using machine vision has significantly improved, particularly when handling heterogeneous textual data. Traditional OCR models struggle with script variations, writing styles, and degraded documents. Advancements in technology are leading to new AI models with improved architecture for handling multiple languages and complex data formats. Despite this progress, a comprehensive evaluation of OCR advancements remains limited. Based on the established preferred reporting items for systematic reviews and meta-analysis (PRISMA) guidelines, this literature review presents an extensive assessment of OCR research to trace the evolution of AI models over the past decade. It explores the transition in AI models, application domains, data types, linguistic coverage, and challenges. Through a detailed analysis of 97 selected studies published during January 2015 - January 2025, key OCR models are identified, and their performance, strengths, and limitations are analyzed. The findings highlight how OCR technologies have evolved to address structured and unstructured text, scene text recognition, and multilingual processing. Unresolved challenges include limited resources for underrepresented languages, high variability in handwritten text, visual similarity among characters, and constraints in real-time OCR applications. To address these issues, several promising approaches are proposed. Key suggestions include self-supervised learning, multimodal AI, automated machine learning (AutoML), AI-assisted postprocessing, tiny machine learning (TinyML), and the creation of joint corpora for script matching. The future recommendations aim to enhance OCR accuracy and tackle the challenges identified for real-time industrial applications. This study will guide future research and establish a foundation for OCR field.

OCR机器学习多语言文献综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。