arXiv:2511.19118cs.CL2025-11

用符号规则统一纳瓦特尔语拼写,提升文本一致性

A symbolic Perl algorithm for the unification of Nahuatl word spellings

  • 基于符号正则表达式实现拼写统一
  • 在π-yalli语料库上验证,多数句子语义保持一致
  • 适合语言学研究与濒危语言数字化

本文描述了一种用于自动统一名为Nawatl的文本拼写的符号模型。该模型基于我们此前用于分析纳瓦特尔语句子的算法,并利用名为$π$-yalli的多拼写语料库。我们的自动统一算法通过符号正则表达式实现语言学规则。同时,我们提出并实施了一种人工评估协议,通过句子语义任务测试生成统一句子的质量。评估者对大多数期望特征表现出积极评价,结果令人鼓舞。

原文摘要 · Abstract (English)

In this paper, we describe a symbolic model for the automatic orthographic unification of Nawatl text documents. Our model is based on algorithms that we have previously used to analyze sentences in Nawatl, and on the corpus called $π$-yalli, consisting of texts in several Nawatl orthographies. Our automatic unification algorithm implements linguistic rules in symbolic regular expressions. We also present a manual evaluation protocol that we have proposed and implemented to assess the quality of the unified sentences generated by our algorithm, by testing in a sentence semantic task. We have obtained encouraging results from the evaluators for most of the desired features of our artificially unified sentences

自然语言处理濒危语言符号计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。