轻量级手写印刷体文本分割框架,兼顾精度与速度。
Handwritten and Printed Text Segmentation via Region-Aware Human-Writing Descriptor Engineering

- 设计区域感知手写特征描述符,捕捉句子级手写变化
- 在MAD-HPTS上比顶尖模型低1.4%准确率但推理快8倍
- 适合资源受限设备部署,对分类器选择不敏感
随着教育和办公场景中纸质文档重用需求增加,手写与印刷体文本的精确分割成为文档数字化的关键步骤。尽管已有多种深度学习模型,但其高计算成本限制了在资源受限边缘设备上的部署。为此,本文提出一种轻量级框架,专为计算能力极弱的设备优化。方法从句级连通域分割算法出发,提取连贯的句子级片段;设计新型区域感知手写特征描述符(RHD),捕获句子级别的人类手写内在变异性;随后可无缝集成简单传统分类器,实现优异的分类性能,表明该描述符对分类器选择具有鲁棒性。在自建多语言高质量标注数据集MAD-HPTS和公开基准PHD-AS上进行大量实验,结果表明本框架在准确率与计算效率上均优于现有最先进方法。在MAD-HPTS上,仅损失1.4%准确率,推理速度提升超8倍,非常适合轻量化部署。
原文摘要 · Abstract (English)
With the increasing demand for reusing paper documents in educational and office settings, accurate segmentation of handwritten and printed text has become a crucial step in document digitization. Although numerous deep learning models have been developed for this task, their high computational cost limits deployment on resource-constrained edge devices. To address this challenge, we present a lightweight framework optimized for efficient performance on devices with severely limited computational capacity. Our approach begins with the Sentence-level Connected Component Segmentation algorithm, aimed at extracting coherent sentence-level segments from document images. We then design a novel Region-aware Handwriting Descriptor (RHD) to capture the intrinsic variability of human handwriting at the sentence level. A simple conventional classifier can then be seamlessly integrated with our designed descriptor, demonstrating strong classification performance for distinguishing handwritten and printed sentence-level text images, highlighting that the proposed descriptor is agnostic to the choice of classifier. Extensive experiments are performed on our self-constructed Multilingual High-Quality Annotated Dataset for Handwritten and Printed Text Segmentation (MAD-HPTS) and a public benchmark PHD-AS, and the experimental results demonstrate that the proposed framework outperforms current state-of-the-art methods in both accuracy and computational efficiency. On MAD-HPTS, our method sacrifices only 1.4% accuracy compared to the leading deep neural network baseline, yet achieves more than 8 times speedup in inference, making it well-suited for lightweight deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。