arXiv:2511.09812cs.CL2025-11

解决高难拼写纠错问题,提升柬埔寨语识别准确率至94.4%

Khmer Spellchecking: A Holistic Approach

  • 融合子词分段、命名实体识别等技术构建综合纠错系统
  • 在真实数据上达到94.4%的拼写纠错准确率,领先现有方法
  • 适合研究低资源语言处理或东南亚语言技术的开发者

相较于英语等高资源语言,柬埔寨语拼写纠错仍面临诸多挑战:词典与分词模型不匹配、单词存在多种书写形式、复合词松散且常不在词典中,以及专有名词因缺乏命名实体识别模型而被误判为拼写错误。现有方案未能有效应对这些问题。本文提出一种综合方法,整合柬埔寨语子词分段、命名实体识别、字符到发音转换(G2P)及语言模型,以识别潜在修正候选并排序最优结果。实验表明,该方法在拼写纠错任务上达到94.4%的最先进准确率。本研究公开了柬埔寨语拼写纠错与命名实体识别的基准数据集。

原文摘要 · Abstract (English)

Compared to English and other high-resource languages, spellchecking for Khmer remains an unresolved problem due to several challenges. First, there are misalignments between words in the lexicon and the word segmentation model. Second, a Khmer word can be written in different forms. Third, Khmer compound words are often loosely and easily formed, and these compound words are not always found in the lexicon. Fourth, some proper nouns may be flagged as misspellings due to the absence of a Khmer named-entity recognition (NER) model. Unfortunately, existing solutions do not adequately address these challenges. This paper proposes a holistic approach to the Khmer spellchecking problem by integrating Khmer subword segmentation, Khmer NER, Khmer grapheme-to-phoneme (G2P) conversion, and a Khmer language model to tackle these challenges, identify potential correction candidates, and rank the most suitable candidate. Experimental results show that the proposed approach achieves a state-of-the-art Khmer spellchecking accuracy of up to 94.4%, compared to existing solutions. The benchmark datasets for Khmer spellchecking and NER tasks in this study will be made publicly available.

拼写纠错低资源语言命名实体识别柬埔寨语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。