arXiv:2506.09009cs.CL2025-06被引 2

将韩语语法标注与通用标签对齐,提升低资源语言分析效果

UD-KSL Treebank v1.3: A semi-automated framework for aligning XPOS-extracted units with UPOS tags

  • 用半自动框架将词性标注序列与通用标签匹配
  • 新标注2998句议论文,使标注数据量提升30%以上
  • 在数据少时显著提升分词与依存句法分析准确率

本文延续对第二语言韩语的通用依赖标注研究,提出一种半自动化框架,可从XPOS序列中识别形态句法结构,并将其与对应的UPOS类别对齐。同时,通过标注2998个议论文句子,扩展了现有第二语言韩语语料库。为评估对齐效果,我们使用两种NLP工具包,在有无对齐数据的两个数据集上微调了第二语言韩语的形态句法分析模型。结果表明,对齐后的数据不仅提升了各标注层间的一致性,还在有限标注数据条件下显著提高了形态词性标注和依存句法分析的准确率。

原文摘要 · Abstract (English)

The present study extends recent work on Universal Dependencies annotations for second-language (L2) Korean by introducing a semi-automated framework that identifies morphosyntactic constructions from XPOS sequences and aligns those constructions with corresponding UPOS categories. We also broaden the existing L2-Korean corpus by annotating 2,998 new sentences from argumentative essays. To evaluate the impact of XPOS-UPOS alignments, we fine-tune L2-Korean morphosyntactic analysis models on datasets both with and without these alignments, using two NLP toolkits. Our results indicate that the aligned dataset not only improves consistency across annotation layers but also enhances morphosyntactic tagging and dependency-parsing accuracy, particularly in cases of limited annotated data.

自然语言处理韩语标注对齐低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。