综述手写文本识别发展,梳理从单词到文档的演进路径。
Handwritten Text Recognition: A Survey
- 按行级与文档级划分识别任务,统一分析框架
- 覆盖从传统方法到深度学习的模型演进历程
- 适合研究者快速掌握领域脉络与未来方向
手写文本识别(HTR)已成为模式识别与机器学习中的关键领域,广泛应用于历史文献保存、现代数据录入和无障碍解决方案。其核心挑战在于手写风格的高度多样性,导致构建鲁棒识别系统极为困难。本文综述了HTR模型的演进,从早期基于启发式的方法,到如今利用深度学习技术的前沿神经网络模型。研究范围也由最初仅能识别单词层级内容,发展为当前端到端的文档级识别。论文将现有工作分为两大识别层级:(1) 行级以内,涵盖单词与行识别;(2) 行级以上,解决段落与文档级挑战。我们提供了一个统一框架,涵盖研究方法、最新基准测试、关键数据集及文献结果讨论。最后,指出当前紧迫的研究挑战,并提出有前景的未来方向,旨在为研究人员与实践者提供领域推进路线图。
原文摘要 · Abstract (English)
Handwritten Text Recognition (HTR) has become an essential field within pattern recognition and machine learning, with applications spanning historical document preservation to modern data entry and accessibility solutions. The complexity of HTR lies in the high variability of handwriting, which makes it challenging to develop robust recognition systems. This survey examines the evolution of HTR models, tracing their progression from early heuristic-based approaches to contemporary state-of-the-art neural models, which leverage deep learning techniques. The scope of the field has also expanded, with models initially capable of recognizing only word-level content progressing to recent end-to-end document-level approaches. Our paper categorizes existing work into two primary levels of recognition: (1) \emph{up to line-level}, encompassing word and line recognition, and (2) \emph{beyond line-level}, addressing paragraph- and document-level challenges. We provide a unified framework that examines research methodologies, recent advances in benchmarking, key datasets in the field, and a discussion of the results reported in the literature. Finally, we identify pressing research challenges and outline promising future directions, aiming to equip researchers and practitioners with a roadmap for advancing the field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。