复杂语言因分词效率低,导致算力成本翻倍、准确率下降。
The Token Tax: Systematic Bias in Multilingual Tokenization
- 用分词数/词比(肥沃度)衡量语言复杂性,发现越高的肥沃度对应越低的准确率。
- 在16种非洲语言上,分词数每翻倍,训练成本和时间均增加为四倍。
- 推理型模型在高低资源语言中表现更好,缩小了以往的性能差距。
分词效率低下对形态复杂的低资源语言造成结构性劣势,增加计算开销并降低准确性。我们在包含9000道多选题、5个主题、16种非洲语言的AfriMMLU数据集上评估了10个大语言模型,发现分词数/词比(肥沃度)能可靠预测准确率:肥沃度越高,准确率越低,且在所有模型和主题中均一致。此外,推理型模型(如DeepSeek, o1)在高/低资源语言上均优于非推理型模型,缩小了以往模型间的性能差距。进一步将分词膨胀转化为经济成本,发现分词数翻倍会导致训练成本与时间变为四倍,凸显诸多语言面临的‘分词税’问题。这些结果推动了形态感知分词、公平定价及多语言基准的发展,以实现更公平的自然语言处理。
原文摘要 · Abstract (English)
Tokenization inefficiency imposes structural disadvantages on morphologically complex, low-resource languages, inflating compute resources and depressing accuracy. We evaluate 10 large language models (LLMs) on AfriMMLU (9,000 MCQA items; 5 subjects; 16 African languages) and show that fertility (tokens/word) reliably predicts accuracy. Higher fertility consistently predicts lower accuracy across all models and subjects. We further find that reasoning models (DeepSeek, o1) consistently outperform non-reasoning peers across high and low resource languages in the AfriMMLU dataset, narrowing accuracy gaps observed in prior generations. Finally, translating token inflation to economics, a doubling in tokens results in quadrupled training cost and time, underscoring the token tax faced by many languages. These results motivate morphologically aware tokenization, fair pricing, and multilingual benchmarks for equitable natural language processing (NLP).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。