CHURRO模型专攻历史文本识别,准确率高且成本低。
CHURRO: Making History Readable with an Open-Weight Large Vision-Language Model for High-Accuracy, Low-Cost Historical Text Recognition
- 用155个历史文集训练30亿参数模型,统一处理多语言、多形态历史文本
- 在测试集上印刷体识别率达82.3%,手写体达70.1%,超越现有最优模型
- 开源模型与数据集,助力文化遗产数字化研究
准确识别历史文献对文化传承与学术研究至关重要。现有视觉语言模型(VLM)针对现代标准文本设计,难以应对历史文本中的多语言、多字体、版式不规则及严重退化问题。本文提出CHURRO,一个30亿参数的开源视觉语言模型,专用于历史文本识别。模型在迄今最大的历史文本识别数据集CHURRO-DS上训练,该数据集整合了155个历史文献集合,共99,491页,覆盖22个世纪的46个语言族系,包括历史变体和已灭绝语言。我们在CHURRO-DS上评估多个开源与闭源VLM及光学字符识别(OCR)系统,发现CHURRO优于所有其他VLM。在测试集上,其印刷体和手写体的归一化莱文斯坦相似度分别达到82.3%和70.1%,较第二名Gemini 2.5 Pro提升1.4%和6.5%,同时成本低15.5倍。通过开放模型与数据集,我们旨在推动社区协作,提升历史文本可读性,加速学术研究。
原文摘要 · Abstract (English)
Accurate text recognition for historical documents can greatly advance the study and preservation of cultural heritage. Existing vision-language models (VLMs), however, are designed for modern, standardized texts and are not equipped to read the diverse languages and scripts, irregular layouts, and frequent degradation found in historical materials. This paper presents CHURRO, a 3B-parameter open-weight VLM specialized for historical text recognition. The model is trained on CHURRO-DS, the largest historical text recognition dataset to date. CHURRO-DS unifies 155 historical corpora comprising 99,491 pages, spanning 22 centuries of textual heritage across 46 language clusters, including historical variants and dead languages. We evaluate several open-weight and closed VLMs and optical character recognition (OCR) systems on CHURRO-DS and find that CHURRO outperforms all other VLMs. On the CHURRO-DS test set, CHURRO achieves 82.3% (printed) and 70.1% (handwritten) normalized Levenshtein similarity, surpassing the second-best model, Gemini 2.5 Pro, by 1.4% and 6.5%, respectively, while being 15.5 times more cost-effective. By releasing the model and dataset, we aim to enable community-driven research to improve the readability of historical texts and accelerate scholarship.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。