arXiv:2410.24034cs.CV2024-10被引 6

用多模态大模型提升古籍手写体识别准确率

Handwriting Recognition in Historical Documents with Multimodal LLM

  • 用Gemini等多模态大模型进行少样本手写体识别
  • 在多个古籍数据集上超越传统Transformer模型
  • 适合古籍数字化与文化遗产保护研究者

大量历史文化遗产仅以手写稿形式存在。相较于印刷体,跨字体、跨书写风格的手写体光学字符识别(OCR)仍极具挑战。尽管基于Transformer的模型表现良好,但依赖大量人工标注数据且泛化能力差。多模态大模型如Gemini在少样本提示下已展现强大视觉与文本理解能力。本文评估Gemini在手写文档转录任务上的表现,结果表明其在多个历史文献数据集上优于当前主流Transformer模型,为大规模古籍数字化提供新路径。

原文摘要 · Abstract (English)

There is an immense quantity of historical and cultural documentation that exists only as handwritten manuscripts. At the same time, performing OCR across scripts and different handwriting styles has proven to be an enormously difficult problem relative to the process of digitizing print. While recent Transformer based models have achieved relatively strong performance, they rely heavily on manually transcribed training data and have difficulty generalizing across writers. Multimodal LLM, such as GPT-4v and Gemini, have demonstrated effectiveness in performing OCR and computer vision tasks with few shot prompting. In this paper, I evaluate the accuracy of handwritten document transcriptions generated by Gemini against the current state of the art Transformer based methods. Keywords: Optical Character Recognition, Multimodal Language Models, Cultural Preservation, Mass digitization, Handwriting Recognitio

手写识别多模态大模型古籍数字化OCR

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。