用三种方法在普通电脑上试了西班牙古籍文字识别,效果尚可。
Transcribing Spanish Texts from the Past: Experiments with Transkribus, Tesseract and Granite
- 对比网页OCR、传统引擎和轻量多模态模型三种方案
- 在消费级硬件上实现可接受的识别准确率
- 适合对古籍数字化感兴趣的中文研究者参考
本文介绍了GRESEL团队在IberLEF 2025共享任务PastReader:转录历史文本中的实验与结果。为参与任务并比较不同方法,开展了三类实验:使用基于网页的OCR服务、传统OCR引擎以及紧凑型多模态模型。所有实验均在消费级硬件上运行,尽管缺乏高性能计算能力,但具备足够存储与稳定性。结果虽令人满意,但仍存改进空间。未来工作将结合BNE提供的西班牙语历史文本数据集,探索新方法与思路。
原文摘要 · Abstract (English)
This article presents the experiments and results obtained by the GRESEL team in the IberLEF 2025 shared task PastReader: Transcribing Texts from the Past. Three types of experiments were conducted with the dual aim of participating in the task and enabling comparisons across different approaches. These included the use of a web-based OCR service, a traditional OCR engine, and a compact multimodal model. All experiments were run on consumer-grade hardware, which, despite lacking high-performance computing capacity, provided sufficient storage and stability. The results, while satisfactory, leave room for further improvement. Future work will focus on exploring new techniques and ideas using the Spanish-language dataset provided by the shared task, in collaboration with Biblioteca Nacional de España (BNE).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。