25种欧洲语言分词器效率差异显著,乌克兰语代价最高。
The Tokenizer Tax Across 25 European Languages: Domain Invariance, Cross-Lingual Few-Shot Effects, and the Ukrainian Penalty
- 测量25种语言分词器的词元密度,发现英语最高效(1.2词元/词)
- 乌克兰语分词代价比同类斯拉夫语言高15-18%,因训练数据不足
- 分词器性能跨文本领域稳定,且少样本迁移能力与模型相关
分词器词元密度(每词词元数)给非英语自然语言处理带来隐性成本。我们在平行文本上对十种基础模型在25种欧洲语言中测量词元密度,首次绘制出该大陆的分词器税地图。词元密度范围从英语的1.2到希腊语/马耳他语的约3.1,呈现清晰层级:罗曼语族(1.5-1.7)、日耳曼语族(1.7-1.9)、斯拉夫语族(2.2-2.5)、乌拉尔/波罗的海语族(2.7-3.0)。乌克兰语(2.7)比同族斯拉夫语言高出15-18%,反映其预训练数据不足。词元密度排名在三种文本类型间高度一致(ρ > 0.97)。子词分析显示,高密度分词器更倾向于破坏形态边界而非保留。在四种斯拉夫语言上的跨语言少样本评估表明,少样本效应是模型内在属性,非语言相关。所有测量结果已公开发布为数据集。
原文摘要 · Abstract (English)
Tokenizer fertility the number of tokens per word imposes a hidden cost on non-English NLP. We measure fertility for ten foundation models across 25 European languages on parallel text, producing the first controlled tokenizer tax map for the continent. The tax spans 2.5x from English (1.2 tokens/word) to Greek/Maltese (~3.1), following a clear hierarchy: Romance (1.5-1.7), Germanic (1.7-1.9), Slavic (2.2-2.5), Uralic/Baltic (2.7-3.0). Ukrainian (2.7) pays 15-18% more than cognate Slavic languages, reflecting underrepresentation in pre-training data. Fertility rankings are domain-invariant across three text registers (rho > 0.97). A subword analysis reveals that high-fertility tokenizers fragment morphological boundaries rather than preserving them. Cross-lingual few-shot evaluation on four Slavic languages shows that few-shot effects are model-intrinsic, not language-dependent. We release all measurements as a public dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。