arXiv:2410.06944cs.CL2024-10EMNLP

提出对比自监督学习,提升低资源富形态语言依存句法分析对词序变化的鲁棒性。

CSSL: Contrastive Self-Supervised Learning for Dependency Parsing on Relatively Free Word Ordered and Morphologically Rich Low Resource Languages

  • 设计对比自监督学习,利用富形态语言自由词序特性增强模型鲁棒性。
  • 在7种自由词序语言上,平均提升UAS 3.03、LAS 2.95分。
  • 适合研究低资源语言句法分析与自监督学习的学者参考。

神经依存句法分析在低资源富形态语言上已取得显著进展。已知富形态语言具有相对自由的词序特征,这引发一个根本问题:能否利用这种自由词序特性,提升依存句法分析性能,使模型对词序变化更具鲁棒性?本文针对7种自由词序语言,评估了基于图的解析架构的鲁棒性,并深入考察数据增强和移除位置编码等关键修改对适配这些架构的影响。为此,我们提出一种对比自监督学习方法,以增强模型对词序变化的适应能力。实验表明,相较于最佳基线模型,该方法在7种语言上平均提升UAS 3.03分、LAS 2.95分。

原文摘要 · Abstract (English)

Neural dependency parsing has achieved remarkable performance for low resource morphologically rich languages. It has also been well-studied that morphologically rich languages exhibit relatively free word order. This prompts a fundamental investigation: Is there a way to enhance dependency parsing performance, making the model robust to word order variations utilizing the relatively free word order nature of morphologically rich languages? In this work, we examine the robustness of graph-based parsing architectures on 7 relatively free word order languages. We focus on scrutinizing essential modifications such as data augmentation and the removal of position encoding required to adapt these architectures accordingly. To this end, we propose a contrastive self-supervised learning method to make the model robust to word order variations. Furthermore, our proposed modification demonstrates a substantial average gain of 3.03/2.95 points in 7 relatively free word order languages, as measured by the UAS/LAS Score metric when compared to the best performing baseline.

依存句法分析自监督学习低资源语言对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。