用最优传输对齐图文,少样本也能提升手写文本识别效果
Optimal Transport for Handwritten Text Recognition in a Low-Resource Regime
- 通过最优传输将无标签图像与语义词向量对齐
- 仅需少量标注数据,迭代生成伪标签提升识别准确率
- 适合历史档案等标注稀缺的场景
手写文本识别(HTR)在文档图像理解中至关重要。现有先进方法依赖大量标注数据,难以应用于历史档案或小规模现代语料等低资源场景。本文提出一种新框架,不依赖传统HTR范式,可利用词汇特征的弱先验知识。该方法采用迭代自举策略,通过最优传输(OT)将无标签图像的视觉特征与语义词表示对齐。从极小标注集出发,框架逐步匹配图像与文本标签,为高置信度对齐生成伪标签,并在不断扩大的数据集上重训练识别器。数值实验表明,该迭代视觉-语义对齐方案在低资源HTR基准上显著提升识别准确率。
原文摘要 · Abstract (English)
Handwritten Text Recognition (HTR) is a task of central importance in the field of document image understanding. State-of-the-art methods for HTR require the use of extensive annotated sets for training, making them impractical for low-resource domains like historical archives or limited-size modern collections. This paper introduces a novel framework that, unlike the standard HTR model paradigm, can leverage mild prior knowledge of lexical characteristics; this is ideal for scenarios where labeled data are scarce. We propose an iterative bootstrapping approach that aligns visual features extracted from unlabeled images with semantic word representations using Optimal Transport (OT). Starting with a minimal set of labeled examples, the framework iteratively matches word images to text labels, generates pseudo-labels for high-confidence alignments, and retrains the recognizer on the growing dataset. Numerical experiments demonstrate that our iterative visual-semantic alignment scheme significantly improves recognition accuracy on low-resource HTR benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。