arXiv:2410.02283cs.CLcs.AI2024-10

评估BETO模型子词词汇的形态质量,发现其与真实词素对齐差。

Morphological evaluation of subwords vocabulary used by BETO language model

  • 用重叠度、凝聚力等三指标量化子词与词素匹配程度
  • BETO子词词汇形态质量低,增大训练语料也无改善
  • 揭露其实际使用Wordpiece算法,与宣称不符

大型语言模型使用的子词分词算法虽高效且可自动构建词汇,但子词未必与真实词素对齐,可能影响模型性能,具体何时发生尚不明确。此前研究提出基于相关性、凝聚力和形态准确率的评估方法,应用于BPE、Wordpiece和Unigram三种算法,发现其词汇普遍形态质量较低。本文将该方法用于BETO模型(基于大规模西班牙语语料训练的BERT模型)的分词器,结果表明其词汇形态质量同样偏低;且增加训练语料规模未能提升其形态质量。此外,评估还揭示其实际采用Wordpiece算法,与作者声称存在不一致。

原文摘要 · Abstract (English)

Subword tokenization algorithms used by Large Language Models are significantly more efficient and can independently build the necessary vocabulary of words and subwords without human intervention. However, those subwords do not always align with real morphemes, potentially impacting the models' performance, though it remains uncertain when this might occur. In previous research, we proposed a method to assess the morphological quality of vocabularies, focusing on the overlap between these vocabularies and the morphemes of a given language. Our evaluation method was built on three quality measures, relevance, cohesion, and morphological accuracy, and a procedure for their assessment. By applying this method to vocabularies created by three subword tokenization algorithms, BPE, Wordpiece, and Unigram, we concluded that these vocabularies generally exhibit very low morphological quality. In this article, we apply this evaluation to the tokenizer of BETO, a BERT language model trained on large Spanish corpora. This evaluation, along with our previous results, helped us conclude that its vocabulary has a low morphological quality, and we also found that training the tokenizer in a larger corpus does not improve the morphological quality of the generated vocabulary. Additionally, this evaluation helps clarify the algorithm used by the tokenizer, that is, Wordpiece, given the inconsistencies between the authors' claims and the model's configuration.

语言模型子词分词形态学BETO

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。