arXiv:2510.12389cs.CLcs.AI2025-10被引 2

非拉丁语系语言在大模型中token化效率更低,导致计算成本更高。

Tokenization Disparities as Infrastructure Bias: How Subword Systems Create Inequities in LLM Access and Efficiency

  • 统一用tiktoken对200多种语言进行分词评估
  • 非拉丁语系语言的分词成本是英文的3-5倍
  • 呼吁构建更包容的多语言分词机制

分词差异严重影响了全球语言群体公平获取人工智能的机会。本研究对200多种语言进行了大规模跨语言分词效率评估,系统量化了大语言模型(LLMs)中的计算不平等。采用标准化实验框架,所有语言样本均经统一预处理和归一化,使用tiktoken库进行一致分词。通过Tokens Per Sentence(TPS)和相对分词成本(RTC)等指标与英语基准对比,发现显著且系统性的差异:拉丁字母语言分词效率更高,而非拉丁语系及形态复杂语言的分词成本普遍高出3-5倍。这些低效导致低资源语言面临更高的计算开销和更差的上下文利用率。研究揭示当前AI系统存在结构性不公,未来应发展考虑语言类型多样性的分词策略与自适应词汇构建方法,推动更具包容性和计算公平性的多语言AI体系。

原文摘要 · Abstract (English)

Tokenization disparities pose a significant barrier to achieving equitable access to artificial intelligence across linguistically diverse populations. This study conducts a large-scale cross-linguistic evaluation of tokenization efficiency in over 200 languages to systematically quantify computational inequities in large language models (LLMs). Using a standardized experimental framework, we applied consistent preprocessing and normalization protocols, followed by uniform tokenization through the tiktoken library across all language samples. Comprehensive tokenization statistics were collected using established evaluation metrics, including Tokens Per Sentence (TPS) and Relative Tokenization Cost (RTC), benchmarked against English baselines. Our cross-linguistic analysis reveals substantial and systematic disparities: Latin-script languages consistently exhibit higher tokenization efficiency, while non-Latin and morphologically complex languages incur significantly greater token inflation, often 3-5 times higher RTC ratios. These inefficiencies translate into increased computational costs and reduced effective context utilization for underrepresented languages. Overall, the findings highlight structural inequities in current AI systems, where speakers of low-resource and non-Latin languages face disproportionate computational disadvantages. Future research should prioritize the development of linguistically informed tokenization strategies and adaptive vocabulary construction methods that incorporate typological diversity, ensuring more inclusive and computationally equitable multilingual AI systems.

分词不平等多语言计算效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。