arXiv:2510.07931cs.CL2025-10被引 1

用大模型自动补全17-18世纪爱沙尼亚-德语词典,提升古籍数字化效率。

Vision-Enabled LLMs in Historical Lexicography: Digitising and Enriching Estonian-German Dictionaries from the 17th and 18th Centuries

  • 结合视觉与语言模型,识别哥特体印刷文字并结构化输出。
  • 在81%的词条中准确生成现代词义和对应词。
  • 适合从事历史语言学、古籍数字化的研究者参考。

本文介绍2022至2025年间爱沙尼亚语言研究所开展的研究,将大语言模型(LLMs)应用于17至18世纪爱沙尼亚语词典的分析。研究聚焦三个方向:以现代词形和意义丰富历史词典;利用视觉增强型LLM对哥特体(Fraktur)印刷文本进行文字识别;为构建跨来源统一数据集做准备。对J. Gutslaff 1648年词典的初步实验表明,当提供足够上下文时,Claude 3.7 Sonnet可准确为81%的词头提供现代释义与对应词。在对A. T. Helle 1732年词典的文字识别实验中,零样本方法成功将41%的词头识别并结构化为无错JSON格式输出。针对A. W. Hupel 1780年语法书中的爱沙尼亚-德语词典部分,采用扫描图像重叠分块策略,由一个LLM负责文字识别,另一个负责整合结构化结果。研究显示,即使对小语种,LLM也能显著节省时间和成本。

原文摘要 · Abstract (English)

This article presents research conducted at the Institute of the Estonian Language between 2022 and 2025 on the application of large language models (LLMs) to the study of 17th and 18th century Estonian dictionaries. The authors address three main areas: enriching historical dictionaries with modern word forms and meanings; using vision-enabled LLMs to perform text recognition on sources printed in Gothic script (Fraktur); and preparing for the creation of a unified, cross-source dataset. Initial experiments with J. Gutslaff's 1648 dictionary indicate that LLMs have significant potential for semi-automatic enrichment of dictionary information. When provided with sufficient context, Claude 3.7 Sonnet accurately provided meanings and modern equivalents for 81% of headword entries. In a text recognition experiment with A. T. Helle's 1732 dictionary, a zero-shot method successfully identified and structured 41% of headword entries into error-free JSON-formatted output. For digitising the Estonian-German dictionary section of A. W. Hupel's 1780 grammar, overlapping tiling of scanned image files is employed, with one LLM being used for text recognition and a second for merging the structured output. These findings demonstrate that even for minor languages LLMs have a significant potential for saving time and financial resources.

历史语言学古籍数字化大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。