为维吾尔语构建了首个语法驱动的依存句法树库,助力语言教学智能化
Construction and educational application of a linguistically grounded dependency treebank for Uyghur
- 基于语言学原则设计标注框架,精准捕捉维吾尔语黏着语特征
- 树库含3456句,交叉弧率降至0.06%,显著优于通用依存标准
- 可直接用于智能语法辅导系统,提升二语学习者理解效率
针对低资源黏着语维吾尔语,现有标注框架与语法结构不匹配的问题,本文提出现代维吾尔语依存树库(MUDT),专门针对其零系词结构和细粒度格标记等复杂特征。采用大模型预标注结合人工校对的混合流程,构建了包含3,456个句子的高质量树库。内在结构评估显示,MUDT将交叉弧率从通用依存标准的7.35%降低至0.06%,显著提升依存投影性。外在解析实验表明,基于MUDT训练的模型在域内准确率和跨域泛化能力上均优于基于UD的基线模型。为进一步验证实用性,开发了基于MUDT的智能语法辅导系统,将句法分析转化为可解释的教学反馈。对35名二语学习者的控制实验表明,获得句法感知反馈的学生学习成效显著高于对照组。研究证实MUDT是句法分析的可靠基础,凸显语言学驱动NLP资源在弥合计算模型与学习者认知需求之间差距的关键作用。
原文摘要 · Abstract (English)
Developing effective educational technologies for low-resource agglutinative languages like Uyghur is often hindered by the mismatch between existing annotation frameworks and specific grammatical structures. To address this challenge, this study introduces the Modern Uyghur Dependency Treebank (MUDT), a linguistically grounded annotation framework specifically designed to capture the agglutinative complexity of Uyghur, including zero copula constructions and fine-grained case marking. Utilizing a hybrid pipeline that combines Large Language Model pre-annotation with rigorous human correction, a high-quality treebank consisting of 3,456 sentences was constructed. Intrinsic structural evaluation reveals that MUDT significantly improves dependency projectivity by reducing the crossing-arc rate from 7.35\% in the Universal Dependencies standard to 0.06\%. Extrinsic parsing experiments using UDPipe and Stanza further demonstrate that models trained on MUDT achieve superior in-domain accuracy and cross-domain generalization compared to UD-based baselines. To validate the practical utility of this computational resource, an AI-assisted grammar tutoring system was developed to translate MUDT-based syntactic analyses into interpretable pedagogical feedback. A controlled experiment involving 35 second-language learners indicated that students receiving syntax-aware feedback achieved significantly higher learning gains compared to those in a control group. These findings establish MUDT as a robust foundation for syntactic analysis and underscore the critical role of linguistically informed natural language processing resources in bridging the gap between computational models and the cognitive needs of second-language learners.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。