用合成数据训练神经拼写检查器,显著提升斯洛文尼亚语纠错效果
Neural spell-checker: Beyond words with synthetic data generation
- 基于大规模合成错误数据训练语言模型进行拼写纠错
- 在斯洛文尼亚语数据集上实现更高准确率与召回率
- 适合自然语言处理研究者及多语言文本工具开发者
拼写检查器能有效提升书面沟通质量,识别文本中的错别字。近年来,深度学习特别是大语言模型的发展,为传统拼写检查器带来新功能,不仅能判断拼写是否正确,还能评估词语在上下文中的适用性。本文提出并比较了两种新型拼写检查器,在合成数据、学习者文本和通用领域斯洛文尼亚语数据集上进行了评估。第一种是基于形态词典的快速词级方法,词表规模远超现有系统;第二种是基于大规模语料库并注入合成错误训练的语言模型。我们详细阐述了训练数据构建策略,发现其对神经拼写检查器性能至关重要。实验表明,所提神经模型在斯洛文尼亚语上的精度和召回率均显著优于所有现有拼写检查器。
原文摘要 · Abstract (English)
Spell-checkers are valuable tools that enhance communication by identifying misspelled words in written texts. Recent improvements in deep learning, and in particular in large language models, have opened new opportunities to improve traditional spell-checkers with new functionalities that not only assess spelling correctness but also the suitability of a word for a given context. In our work, we present and compare two new spell-checkers and evaluate them on synthetic, learner, and more general-domain Slovene datasets. The first spell-checker is a traditional, fast, word-based approach, based on a morphological lexicon with a significantly larger word list compared to existing spell-checkers. The second approach uses a language model trained on a large corpus with synthetically inserted errors. We present the training data construction strategies, which turn out to be a crucial component of neural spell-checkers. Further, the proposed neural model significantly outperforms all existing spell-checkers for Slovene in both precision and recall.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。