arXiv:2411.12240cs.CLcs.AI2024-11被引 11

对比12个大模型在22种印度官方语言上的分词性能,发现SUTRA表现最优。

Evaluating Tokenizer Performance of Large Language Models Across Official Indian Languages

  • 用归一化序列长度评估12个模型在22种印地语系语言的分词效率
  • SUTRA tokenizer在14种语言中表现最佳,优于多个专用模型
  • GPT-4o比GPT-4更擅长处理印度语言,而Project Indus表现有限

基于Transformer架构的大语言模型(LLMs)在多个领域引发变革,其中分词在预处理与微调阶段起关键作用。针对印地语系多语言模型,有效的分词对性能优化至关重要。本文全面评估了12个大模型在印度全部22种官方语言上的分词表现,重点比较其分词效率。采用归一化序列长度(NSL)作为核心指标。结果表明,SUTRA分词器在14种语言中优于其他所有模型,包括多个专为印地语设计的模型。重要发现包括:SUTRA分词器对印地语系语言具有显著优势;GPT-4o相比前代模型GPT-4在印度语言处理上实现进步;而Project Indus在部分语言中表现不佳。本研究强调为多语言及印地语中心模型开发针对性分词策略的重要性,为未来分词器设计提升语言覆盖与模型效率奠定基础。

原文摘要 · Abstract (English)

Large Language Models (LLMs) based on transformer architectures have revolutionized a variety of domains, with tokenization playing a pivotal role in their pre-processing and fine-tuning stages. In multilingual models, particularly those tailored for Indic languages, effective tokenization is crucial for optimizing performance. This paper presents a comprehensive evaluation of tokenizers used by 12 LLMs across all 22 official languages of India, with a focus on comparing the efficiency of their tokenization processes. We employed the Normalized Sequence Length (NSL) as a key metric in our analysis. Our findings reveal that the SUTRA tokenizer outperforms all other models, including several Indic-specific models, excelling in 14 languages. Notable insights include the SUTRA tokenizer's superior handling of Indic languages, GPT-4o's advancement over its predecessor GPT-4 in processing Indian languages, and the limited performance of Project Indus in certain languages. This study underscores the critical importance of developing targeted tokenization strategies for multilingual and Indic-centric models, laying the groundwork for future improvements in tokenizer design to enhance linguistic coverage and model efficiency.

分词器大模型印地语多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。