Qtok工具评估多语言分词器质量,发现现有分词存在语言偏见。
Qtok: A Comprehensive Framework for Evaluating Multilingual Tokenizer Quality in Large Language Models
- 提出语言覆盖、标记完整性和分布等多维度评估指标
- 分析13个分词器在58个模型中的表现,发现语言分布差异大
- 适合关注多语言模型分词效果的研究者使用
在大型语言模型(LLM)开发中,训练数据集质量备受关注,但分词器的作用,尤其在多语言场景下的影响却较少被重视。分词质量会显著影响模型对多种语言的处理能力。本文提出Qtok,一个专注于多语言环境下分词器质量评估的工具。研究设计了一套评估指标,包括语言覆盖度、标记完整性以及跨语言和语言类别分布情况。将这些指标应用于58个公开模型中的13种不同分词器,分析其在各类语言上下文中的输出表现。结果显示,不同语言与语言类别间的标记分布存在显著差异,揭示了当前分词策略中存在的潜在偏见和改进空间。该研究为多语言LLM开发中的分词器评估提供了系统方法,强调了分词环节对多语言能力的关键作用。Qtok工具及分析方法可帮助研究人员评估并优化多语言应用中的分词策略,提供跨指标比较手段,适用于特定多语言任务的分词器选型与调整。
原文摘要 · Abstract (English)
In the development of Large Language Models (LLMs), considerable attention has been given to the quality of training datasets. However, the role of tokenizers in the LLM training pipeline, particularly for multilingual models, has received less focus. The quality of tokenization can significantly impact a model's ability to handle diverse languages effectively. We introduce Qtok, a tool designed to assess tokenizer quality with a specific emphasis on their performance in multilingual contexts. Our research proposes a set of metrics for evaluating tokenizer quality, including measures of language coverage, token completeness, and distribution across languages and linguistic categories. Qtok applies these metrics to evaluate 13 distinct tokenizers from 58 publicly available models, analyzing their output across different linguistic contexts. Our analysis revealed significant variations in token distribution across languages and categories, highlighting potential biases and areas for improvement in current tokenization strategies. This research contributes to the field of tokenizer evaluation within multilingual LLM development by providing a systematic approach to assessing tokenizer quality. Our findings highlight the critical role of tokenization in multilingual LLM capability. The Qtok tool and our analysis methodology offer practical means for researchers to evaluate and improve tokenization strategies for multilingual applications. We offer a method to compare tokenizer quality across these metrics, which may be useful when selecting or adjusting tokenizers for specific multilingual LLM applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。