arXiv:2410.12055cs.CL2024-10被引 2

对比六模型,找到古希腊语分词句法标注最佳方案

A State-of-the-Art Morphosyntactic Parser and Lemmatizer for Ancient Greek

  • 用古希腊语语料训练并比较六种模型的标注能力
  • Trankit在句法标注上最优,GreTa在词形还原上表现最好
  • 强调词向量需配合特定建模策略才能提升效果

本文通过对比六种模型,寻找符合古希腊语依存树库标注规范的先进分词句法分析器与词形还原器。使用标准化的主流标注文本集合,(i) 以随机初始化字符嵌入训练基线模型 Dithrax,(ii) 微调 Trankit 及四种在古希腊语语料上预训练的模型:GreBERTa 与 PhilBERTa 用于分词句法标注,GreTA 与 PhilTa 用于词形还原。贝叶斯分析显示,Dithrax 与 Trankit 在形态标注上几乎等效,句法标注由 Trankit 最优,词形还原则由 GreTA 领先。实验表明,仅靠词嵌入无法获得高 UAS 与 LAS 分数,必须结合专门设计的建模策略以捕捉句法关系。数据集与最优模型已公开可复用。

原文摘要 · Abstract (English)

This paper presents an experiment consisting in the comparison of six models to identify a state-of-the-art morphosyntactic parser and lemmatizer for Ancient Greek capable of annotating according to the Ancient Greek Dependency Treebank annotation scheme. A normalized version of the major collections of annotated texts was used to (i) train the baseline model Dithrax with randomly initialized character embeddings and (ii) fine-tune Trankit and four recent models pretrained on Ancient Greek texts, i.e., GreBERTa and PhilBERTa for morphosyntactic annotation and GreTA and PhilTa for lemmatization. A Bayesian analysis shows that Dithrax and Trankit annotate morphology practically equivalently, while syntax is best annotated by Trankit and lemmata by GreTa. The results of the experiment suggest that token embeddings are not sufficient to achieve high UAS and LAS scores unless they are coupled with a modeling strategy specifically designed to capture syntactic relationships. The dataset and best-performing models are made available online for reuse.

古希腊语句法分析词形还原NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。