提出新指标STRR,揭示多语言分词中英语优先、中文支持强、印地语碎片化的公平性问题
Beyond Fertility: Analyzing STRR as a Metric for Multilingual Tokenization Evaluation
- 引入单标记保留率STRR,衡量词语被保留为单个标记的比例
- 发现中文分词肥大(高保真度),印地语严重碎片化,英语系统性优先
- 适合关注多语言模型公平性的研究者与开发者参考
分词是大语言模型中的关键但评估不足的步骤。传统指标‘肥力’(每词平均标记数)仅反映压缩效率,却掩盖了词汇在不同语言和领域间的分配差异。我们分析了六种常用分词器在七种语言及两个领域中的表现,发现英语肥力稳定,中文肥力较高,且对领域变化不敏感。为弥补肥力的盲点,本文提出单标记保留率(STRR),衡量词语作为单个标记保留的比例。结果表明:英文存在系统性优先,中文获得强力支持,而印地语则出现明显碎片化,提供了可解释的跨语言公平性视角。实验证明STRR能有效补充肥力,为设计更公平的多语言分词器提供实践指导。
原文摘要 · Abstract (English)
Tokenization is a crucial but under-evaluated step in large language models (LLMs). The standard metric, fertility (the average number of tokens per word), captures compression efficiency but obscures how vocabularies are allocated across languages and domains. We analyze six widely used tokenizers across seven languages and two domains, finding stable fertility for English, high fertility for Chinese, and little domain sensitivity. To address fertility's blind spots, we propose the Single Token Retention Rate (STRR), which measures the proportion of words preserved as single tokens. STRR reveals systematic prioritization of English, strong support for Chinese, and fragmentation in Hindi, offering an interpretable view of cross-lingual fairness. Our results show that STRR complements fertility and provides practical guidance for designing more equitable multilingual tokenizers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。