对比五种大模型分词器在阿萨姆语中的表现,发现SUTRA分词器最优。
Performance Evaluation of Tokenizers in Large Language Models for the Assamese Language
- 在阿萨姆语上评估五种主流大模型分词器的性能
- SUTRA分词器平均归一化序列长度仅0.45,表现最佳
- 适合关注低资源语言多语支持的研究者参考
分词器训练对深度学习模型性能至关重要。本研究旨在评估五种前沿大语言模型(LLMs)在印度阿萨姆语中的分词器表现。该研究对理解低资源语言如阿萨姆语的多语言支持能力具有重要意义。结果表明,Two AI开发的SUTRA分词器表现最佳,平均归一化序列长度(NSL)为0.45;紧随其后的是OpenAI的GPT-4o,平均NSL为0.54;Gemina 2、Meta Llama 3.1和Mistral Large Instruct 2407的平均NSL分别为0.82、1.4和1.48。
原文摘要 · Abstract (English)
Training of a tokenizer plays an important role in the performance of deep learning models. This research aims to understand the performance of tokenizers in five state-of-the-art (SOTA) large language models (LLMs) in the Assamese language of India. The research is important to understand the multi-lingual support for a low-resourced language such as Assamese. Our research reveals that the tokenizer of SUTRA from Two AI performs the best with an average Normalized Sequence Length (NSL) value of 0.45, closely followed by the tokenizer of GPT-4o from Open AI with an average NSL value of 0.54, followed by Gemma 2, Meta Llama 3.1, and Mistral Large Instruct 2407 with an average NSL value of 0.82, 1.4, and 1.48 respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。