arXiv:2501.09326cs.CL2025-01

用语义网络解析斯瓦希里语,无需大量训练数据即可实现问答

Algorithm for Semantic Network Generation from Texts of Low Resource Languages Such as Kiswahili

  • 基于主谓宾结构映射文本为语义三元组
  • 在斯瓦希里语问答任务中达到78.6%精确匹配
  • 适合资源匮乏语言的自然语言处理应用

低资源语言(如斯瓦希里语)的机器学习处理因缺乏足够训练数据而困难。但这些语言在日常交流中仍具重要性,用户亟需摘要、消歧与问答等实用处理能力。一种绕过数据依赖的方法是构建语义网络。由于斯瓦希里语具有主谓宾(SVO)结构,而语义网络也是主-谓-宾三元组,因此可将词性标注的SVO成分直接映射为语义三元组。本文提出一种算法,可将原始自然语言文本转换为语义网络,在斯瓦希里语问答任务上达到最高78.6%的精确匹配率。

原文摘要 · Abstract (English)

Processing low-resource languages, such as Kiswahili, using machine learning is difficult due to lack of adequate training data. However, such low-resource languages are still important for human communication and are already in daily use and users need practical machine processing tasks such as summarization, disambiguation and even question answering (QA). One method of processing such languages, while bypassing the need for training data, is the use semantic networks. Some low resource languages, such as Kiswahili, are of the subject-verb-object (SVO) structure, and similarly semantic networks are a triple of subject-predicate-object, hence SVO parts of speech tags can map into a semantic network triple. An algorithm to process raw natural language text and map it into a semantic network is therefore necessary and desirable in structuring low resource languages texts. This algorithm tested on the Kiswahili QA task with upto 78.6% exact match.

语义网络低资源语言斯瓦希里语问答系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。