arXiv:2503.15195cs.CV2025-03被引 18

对比大模型与传统模型手写文字识别效果,发现闭源模型表现更优。

Benchmarking Large Language Models for Handwritten Text Recognition

  • 用大语言模型直接识别手写文本,无需专门训练。
  • 闭源模型如Claude 3.5 Sonnet在零样本下表现最佳,尤其对英文识别优势明显。
  • 模型自我纠错能力弱,且与传统Transkribus系统无明显优劣之分。

传统手写文字识别(HTR)依赖监督学习,需大量人工标注,且布局与文本处理分离导致错误频发。多模态大语言模型(MLLMs)提供无需特定训练即可识别多样手写风格的通用方案。本研究对比了多种私有及开源大模型与Transkribus系统的性能,评估其在英、法、德、意语现代与历史文本上的表现,并测试模型自主修正生成结果的能力。结果表明,在零样本设置下,私有模型(尤其是Claude 3.5 Sonnet)优于开源模型;MLLM在现代手写识别中表现优异,但因预训练数据分布对英语有偏好。与Transkribus相比,两者无一致优势。此外,大模型在零样本转录中自我纠错能力有限。

原文摘要 · Abstract (English)

Traditional machine learning models for Handwritten Text Recognition (HTR) rely on supervised training, requiring extensive manual annotations, and often produce errors due to the separation between layout and text processing. In contrast, Multimodal Large Language Models (MLLMs) offer a general approach to recognizing diverse handwriting styles without the need for model-specific training. The study benchmarks various proprietary and open-source LLMs against Transkribus models, evaluating their performance on both modern and historical datasets written in English, French, German, and Italian. In addition, emphasis is placed on testing the models' ability to autonomously correct previously generated outputs. Findings indicate that proprietary models, especially Claude 3.5 Sonnet, outperform open-source alternatives in zero-shot settings. MLLMs achieve excellent results in recognizing modern handwriting and exhibit a preference for the English language due to their pre-training dataset composition. Comparisons with Transkribus show no consistent advantage for either approach. Moreover, LLMs demonstrate limited ability to autonomously correct errors in zero-shot transcriptions.

手写识别大模型零样本多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。