arXiv:2410.03568cs.CLcs.LG2024-10被引 3

提出兼顾语言特性和多语言支持的分词方法,提升低资源语言模型表现。

Towards Linguistically-Aware and Language-Independent Tokenization for Large Language Models (LLMs)

  • 设计语言感知的子词分词机制,减少对高资源语言的依赖。
  • 在多语言数据上验证分词一致性,降低低资源语言服务成本。
  • 适合关注AI公平性与跨语言应用的研究者和开发者。

本文系统研究了当前先进大语言模型(如GPT-4、GPT-3、DaVinci及BERT base)所采用的分词技术及其对不同语言服务成本与可用性的影响,尤其关注低资源语言。分析涵盖cl100k_base、p50k_base、r50k_base等分词器,揭示子词分词在语言表征上的差异性。研究表明,应推动更具语言意识的开发实践,以提升对传统低资源语言的支持。通过电子健康记录(EHR)系统的案例研究,凸显分词选择的实际影响。本研究旨在促进AI服务在该领域及其他场景中的可推广国际化(I18N)实践,强调包容性,尤其关注在主流AI应用中被忽视的语言。

原文摘要 · Abstract (English)

This paper presents a comprehensive study on the tokenization techniques employed by state-of-the-art large language models (LLMs) and their implications on the cost and availability of services across different languages, especially low resource languages. The analysis considers multiple LLMs, including GPT-4 (using cl100k_base embeddings), GPT-3 (with p50k_base embeddings), and DaVinci (employing r50k_base embeddings), as well as the widely used BERT base tokenizer. The study evaluates the tokenization variability observed across these models and investigates the challenges of linguistic representation in subword tokenization. The research underscores the importance of fostering linguistically-aware development practices, especially for languages that are traditionally under-resourced. Moreover, this paper introduces case studies that highlight the real-world implications of tokenization choices, particularly in the context of electronic health record (EHR) systems. This research aims to promote generalizable Internationalization (I18N) practices in the development of AI services in this domain and beyond, with a strong emphasis on inclusivity, particularly for languages traditionally underrepresented in AI applications.

分词方法多语言LLMAI公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。