arXiv:2411.14198cs.CL2024-11被引 79

语言模型在形态复杂语言上表现差,主因是数据量不足而非语言本身难学。

Why do language models perform worse for morphologically complex languages?

  • 用新指标MorphScore评估分词器对形态结构的匹配程度
  • 数据集规模相当且按字节效率调整后,性能差距大幅缩小
  • 研究结果对提升低资源语言模型性能有重要指导意义

语言模型在不同语言上的表现存在差异。已有研究认为形态类型可能解释部分变异(Cotterell et al., 2018)。我们复现并扩展了相关分析,发现黏着语(如土耳其语)与屈折语(如英语)之间存在性能差距,其中屈折语表现更优。为探究原因,我们提出并检验三个假设:分词器形态对齐、分词质量及数据集大小与测量差异。通过提出新指标MorphScore及22种语言的配套数据集测试形态对齐假设,发现分词质量有一定影响,但形态对齐无显著作用。关键发现是:当不同语言的数据集按‘字节溢价’(byte-premium)标准化后,且规模相当,性能差距显著缩小。这表明语言模型的学习难度并非由形态类型决定,而是源于数据量不均衡。该结果对改善低资源和表现欠佳语言的模型性能具有重要意义。

原文摘要 · Abstract (English)

Language models perform differently across languages. It has been previously suggested that morphological typology may explain some of this variability (Cotterell et al., 2018). We replicate previous analyses and find additional new evidence for a performance gap between agglutinative and fusional languages, where fusional languages, such as English, tend to have better language modeling performance than morphologically more complex languages like Turkish. We then propose and test three possible causes for this performance gap: morphological alignment of tokenizers, tokenization quality, and disparities in dataset sizes and measurement. To test the morphological alignment hypothesis, we present MorphScore, a tokenizer evaluation metric, and supporting datasets for 22 languages. We find some evidence that tokenization quality explains the performance gap, but none for the role of morphological alignment. Instead we find that the performance gap is most reduced when training datasets are of equivalent size across language types, but only when scaled according to the so-called "byte-premium" -- the different encoding efficiencies of different languages and orthographies. These results suggest that no language is harder or easier for a language model to learn on the basis of its morphological typology. Differences in performance can be attributed to disparities in dataset size. These results bear on ongoing efforts to improve performance for low-performing and under-resourced languages.

语言模型形态学数据偏差多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。