arXiv:2409.02841cs.CL2024-09被引 3

用混合模型将1700-1900年德语古籍拼写转为现代标准拼写。

Historical German Text Normalization Using Type- and Token-Based Language Modeling

  • 结合编码器-解码器与因果语言模型,分步完成词型与上下文修正。
  • 在平行语料上达到当前最优准确率,接近大型端到端系统表现。
  • 适合从事历史文本数字化、古籍可读性提升的研究者使用。

历史拼写变异给古籍全文检索与自然语言处理带来挑战。为缩小历史拼写与现代拼写间的差距,通常需对历史文本进行自动正字归一化。本文提出一种针对1700–1900年德语文学文本的归一化系统,基于平行语料训练。该系统采用机器学习方法,结合编码器-解码器模型对单个词型进行归一,再利用预训练因果语言模型在上下文中微调结果。大量实验表明,该系统在准确性上达到当前最优水平,与需微调大规模端到端Transformer语言模型的复杂系统相当。然而,由于模型泛化困难及高质量平行数据稀缺,历史文本归一化仍具挑战。

原文摘要 · Abstract (English)

Historic variations of spelling poses a challenge for full-text search or natural language processing on historical digitized texts. To minimize the gap between the historic orthography and contemporary spelling, usually an automatic orthographic normalization of the historical source material is pursued. This report proposes a normalization system for German literary texts from c. 1700-1900, trained on a parallel corpus. The proposed system makes use of a machine learning approach using Transformer language models, combining an encoder-decoder model to normalize individual word types, and a pre-trained causal language model to adjust these normalizations within their context. An extensive evaluation shows that the proposed system provides state-of-the-art accuracy, comparable with a much larger fully end-to-end sentence-based normalization system, fine-tuning a pre-trained Transformer large language model. However, the normalization of historical text remains a challenge due to difficulties for models to generalize, and the lack of extensive high-quality parallel data.

文本归一化古籍处理Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。