扩充韩语二语依存句法库,提升模型对二语韩语的分析能力。
Second language Korean Universal Dependency treebank v1.2: Focus on data augmentation and annotation scheme refinement
- 新增5454句人工标注数据,优化标注规范
- 微调后模型在域内与域外数据上表现显著提升
- 适合研究二语语言模型与语法分析的学者
我们扩展了第二语言(L2)韩语通用依存句法库,新增5,454句人工标注语料,并修订标注指南以更好地契合通用依存框架。基于此增强后的语料库,我们对三款韩语语言模型进行微调,并在域内与域外的L2韩语数据集上评估其性能。结果表明,微调显著提升了模型在各类指标上的表现,凸显了使用专门针对二语设计的数据集对通用语言模型进行微调,在二语句法分析中的重要性。
原文摘要 · Abstract (English)
We expand the second language (L2) Korean Universal Dependencies (UD) treebank with 5,454 manually annotated sentences. The annotation guidelines are also revised to better align with the UD framework. Using this enhanced treebank, we fine-tune three Korean language models and evaluate their performance on in-domain and out-of-domain L2-Korean datasets. The results show that fine-tuning significantly improves their performance across various metrics, thus highlighting the importance of using well-tailored L2 datasets for fine-tuning first-language-based, general-purpose language models for the morphosyntactic analysis of L2 data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。