用土耳其语树库迁移法,高效构建吉尔吉斯语句法语料
Syntactic Transfer to Kyrgyz Using the Treebank Translation Method
- 通过树库翻译法将土耳其语句法标注迁移到吉尔吉斯语
- 在TueCL语料上准确率优于仅用吉尔吉斯语训练的模型
- 提供标注复杂度评估方法,助力优化人工标注流程
吉尔吉斯语作为低资源语言,构建高质量句法语料需大量人力。本文提出一种基于树库翻译法的工具,将土耳其语的句法标注迁移到吉尔吉斯语,简化语料开发流程。在TueCL树库上的实验表明,该方法相比仅使用吉尔吉斯语KTMU树库训练的单语模型,实现了更高的句法标注准确率。此外,研究还提出一种评估生成句法树人工标注复杂度的方法,有助于进一步优化标注流程。
原文摘要 · Abstract (English)
The Kyrgyz language, as a low-resource language, requires significant effort to create high-quality syntactic corpora. This study proposes an approach to simplify the development process of a syntactic corpus for Kyrgyz. We present a tool for transferring syntactic annotations from Turkish to Kyrgyz based on a treebank translation method. The effectiveness of the proposed tool was evaluated using the TueCL treebank. The results demonstrate that this approach achieves higher syntactic annotation accuracy compared to a monolingual model trained on the Kyrgyz KTMU treebank. Additionally, the study introduces a method for assessing the complexity of manual annotation for the resulting syntactic trees, contributing to further optimization of the annotation process.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。