arXiv:2506.03101cs.CL2025-06ACL被引 11

小模型可低成本预测大模型分词器效果,提升多语言任务选择效率。

Beyond Text Compression: Evaluating Tokenizers Across Scales

  • 用小模型替代大模型评估分词器性能,节省大量计算资源。
  • 多语言场景下分词器差异显著影响模型表现,英语任务则影响不大。
  • 提出基于齐普夫定律的新指标,更准确预测未见语言的分词效果。

分词器的选择会显著影响语言模型性能,但目前缺乏可靠且易获取的分词器评估方法。受缩放一致性启发,我们发现小型模型可在极低计算成本下准确预测分词器对大型模型的影响。系统评估了以英语为中心和多语言分词器后发现:在英语任务中分词器影响几乎可忽略,但在多语言设置下存在一致性能差异。我们提出受齐普夫定律启发的新内在分词器评估指标,其与下游性能的相关性优于文本压缩指标,尤其在建模未见语言时表现更优。通过整合多个指标构建多维度评估框架,为未来语言模型开发中的分词器选择提供了高效可靠的路径。

原文摘要 · Abstract (English)

The choice of tokenizer can profoundly impact language model performance, yet accessible and reliable evaluations of tokenizer quality remain an open challenge. Inspired by scaling consistency, we show that smaller models can accurately predict significant differences in tokenizer impact on larger models at a fraction of the compute cost. By systematically evaluating both English-centric and multilingual tokenizers, we find that tokenizer choice has negligible effects on tasks in English but results in consistent performance differences in multilingual settings. We propose new intrinsic tokenizer metrics inspired by Zipf's law that correlate more strongly with downstream performance than text compression when modeling unseen languages. By combining several metrics to capture multiple aspects of tokenizer behavior, we develop a reliable framework for intrinsic tokenizer evaluations. Our work offers a more efficient path to informed tokenizer selection in future language model development.

分词器评估多语言模型缩放

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。