arXiv:2411.03340cs.CVcs.CL2024-11被引 2

用大模型自动转录历史手稿,准确率超传统工具,速度更快成本更低。

Unlocking the Archives: Using Large Language Models to Transcribe Handwritten Historical Documents

  • 利用多模态大模型批量转录手稿,自动纠错提升精度。
  • 字符错误率低至1.8%,词错误率3.5%,接近人工水平。
  • 适合历史学者、档案馆及数字化项目快速处理海量手稿。

本研究展示大型语言模型(LLMs)在转录历史手稿方面显著优于专用手写文字识别(HTR)软件,且速度更快、成本更低。我们推出了开源工具Transcription Pearl,利用OpenAI、Anthropic和Google的商用多模态LLMs,自动批量转录并校正手稿。在18至19世纪英语手稿的多样化语料库测试中,LLMs的字符错误率(CER)为5.7%至7%,词错误率(WER)为8.9%至15.9%,较顶尖HTR工具Transkribus分别提升14%和32%。更关键的是,当LLMs用于校正自身及传统HTR生成的文本时,达到近似人类水平的准确率:CER低至1.8%,WER为3.5%。整个过程比专有HTR程序快50倍,成本约为其1/50。结果表明,将LLMs集成至Transcription Pearl等工具中,可实现高效、低成本、高精度的历史手稿大规模数字化。

原文摘要 · Abstract (English)

This study demonstrates that Large Language Models (LLMs) can transcribe historical handwritten documents with significantly higher accuracy than specialized Handwritten Text Recognition (HTR) software, while being faster and more cost-effective. We introduce an open-source software tool called Transcription Pearl that leverages these capabilities to automatically transcribe and correct batches of handwritten documents using commercially available multimodal LLMs from OpenAI, Anthropic, and Google. In tests on a diverse corpus of 18th/19th century English language handwritten documents, LLMs achieved Character Error Rates (CER) of 5.7 to 7% and Word Error Rates (WER) of 8.9 to 15.9%, improvements of 14% and 32% respectively over specialized state-of-the-art HTR software like Transkribus. Most significantly, when LLMs were then used to correct those transcriptions as well as texts generated by conventional HTR software, they achieved near-human levels of accuracy, that is CERs as low as 1.8% and WERs of 3.5%. The LLMs also completed these tasks 50 times faster and at approximately 1/50th the cost of proprietary HTR programs. These results demonstrate that when LLMs are incorporated into software tools like Transcription Pearl, they provide an accessible, fast, and highly accurate method for mass transcription of historical handwritten documents, significantly streamlining the digitization process.

手稿转录大模型应用历史数字化多模态模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。