用规则系统解析盖尔语词形变化,助力低资源语言教学工具开发。
A Rule-based Computational Model for Gaidhlig Morphology
- 基于维基词典数据构建可解释的规则引擎,支持词形推导。
- 通过SQL查询提取词汇模式,实现盖尔语词形生成。
- 适合语言教育工具开发者与低资源语言研究者使用。
语言模型和软件工具对维护小语种活力至关重要;然而,当前流行的神经模型需大量训练数据,而这类低资源语言通常缺乏。本文介绍基于维基词典数据构建盖尔语词形规则系统的进展,论证规则方法能有效利用有限样本数据,提升可解释性,并为教学材料设计提供洞见。研究探索了用SQL查询不同词法模式的出现情况,提出一种声明式规则库,配合Python工具可推导盖尔语单词的屈折形式。该功能可用于支持教学工具(如解释语言规律)或更高层级的规则依赖解析器。该方法使维基词典现有数据适用于新场景,增加其应用价值。
原文摘要 · Abstract (English)
Language models and software tools are essential to support the continuing vitality of lesser-used languages; however, currently popular neural models require considerable data for training, which normally is not available for such low-resource languages. This paper describes work-in-progress to construct a rule-based model of Gaidhlig morphology using data from Wiktionary, arguing that rule-based systems effectively leverage limited sample data, support greater interpretability, and provide insights useful in the design of teaching materials. The use of SQL for querying the occurrence of different lexical patterns is investigated, and a declarative rule-base is presented that allows Python utilities to derive inflected forms of Gaidhlig words. This functionality could be used to support educational tools that teach or explain language patterns, for example, or to support higher level tools such as rule-based dependency parsers. This approach adds value to the data already present in Wiktionary by adapting it to new use-cases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。