多模态大模型让古籍文字识别更准更快,无需调参也能自动纠错和提取信息。
Multimodal LLMs for OCR, OCR Post-Correction, and Named Entity Recognition in Historical Documents
- 用多模态大模型直接读图识字,不需预处理或微调。
- 识别错误率低于1%(CER),比传统OCR和同类模型大幅领先。
- 可自动提取人名地名等信息,适合历史文献数字化研究者。
我们研究多模态大语言模型(mLLMs)在历史文献中的应用,涵盖光学字符识别(OCR)、OCR后纠错及命名实体识别(NER)。以1754至1870年间出版的德语城市名录为数据集,对比了现成mLLM与传统OCR模型的转录准确率。结果显示,最优mLLM显著超越主流OCR及其他前沿mLLM。首次提出使用mLLMs对OCR输出进行多模态后纠错,无需图像预处理或模型微调,即可将错误率降至低于1%(CER)。同时,mLLMs能高效识别历史文本中的实体并结构化输出。这些发现初步验证了mLLMs在历史数据采集与文档转录中可能引发范式变革的潜力。
原文摘要 · Abstract (English)
We explore how multimodal Large Language Models (mLLMs) can help researchers transcribe historical documents, extract relevant historical information, and construct datasets from historical sources. Specifically, we investigate the capabilities of mLLMs in performing (1) Optical Character Recognition (OCR), (2) OCR Post-Correction, and (3) Named Entity Recognition (NER) tasks on a set of city directories published in German between 1754 and 1870. First, we benchmark the off-the-shelf transcription accuracy of both mLLMs and conventional OCR models. We find that the best-performing mLLM model significantly outperforms conventional state-of-the-art OCR models and other frontier mLLMs. Second, we are the first to introduce multimodal post-correction of OCR output using mLLMs. We find that this novel approach leads to a drastic improvement in transcription accuracy and consistently produces highly accurate transcriptions (<1% CER), without any image pre-processing or model fine-tuning. Third, we demonstrate that mLLMs can efficiently recognize entities in transcriptions of historical documents and parse them into structured dataset formats. Our findings provide early evidence for the long-term potential of mLLMs to introduce a paradigm shift in the approaches to historical data collection and document transcription.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。