arXiv:2506.23288cs.CL2025-06

用大模型解决历史文献拼写规范化问题,效果优于传统方法。

Two Spelling Normalization Approaches Based on Large Language Models

  • 基于大语言模型设计无监督与机器翻译两种新方法
  • 在多语言跨时期数据集上均取得良好效果
  • 机器翻译类方法更适合该任务,适合人文学者使用

历史文献因缺乏标准化拼写规范且语言自然演化,给人文学科研究带来长期挑战。拼写规范化旨在将文本拼写统一至现代标准。本文提出两种基于大语言模型的新方法:一种为无监督训练,另一种为机器翻译训练。在涵盖多种语言和历史时期的多个数据集上进行评估,结果表明两者均表现良好,但统计机器翻译方法仍是最适配该任务的技术。

原文摘要 · Abstract (English)

The absence of standardized spelling conventions and the organic evolution of human language present an inherent linguistic challenge within historical documents, a longstanding concern for scholars in the humanities. Addressing this issue, spelling normalization endeavors to align a document's orthography with contemporary standards. In this study, we propose two new approaches based on large language models: one of which has been trained without a supervised training, and a second one which has been trained for machine translation. Our evaluation spans multiple datasets encompassing diverse languages and historical periods, leading us to the conclusion that while both of them yielded encouraging results, statistical machine translation still seems to be the most suitable technology for this task.

拼写归一化大模型历史文本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。