arXiv:2412.03160cs.CL2024-12被引 3

揭示词元化本质是字符串的逆同态映射,解释其如何保持语言结构。

Byte BPE Tokenization as an Inverse string Homomorphism

  • 词元化可视为字符串与词元间的逆同态映射。
  • 不同词元化算法均保持源语言的结构特性。
  • 理论分析表明词元化不影响模型对上下文无关语言的识别能力。

词元化是大语言模型训练与推理中的关键预处理步骤。尽管对大语言模型神经架构的表达能力已有广泛研究,但词元化的影响尚未被充分理解。本文证明,无论采用何种算法,词元化在本质上都是字符串与词元之间的逆同态映射。这表明源语言的字符空间与目标语言的词元空间具有同态关系,能够保留源语言的结构特性。此外,我们探讨了‘合理词元化’的概念,即从分词器返回的无歧义词元化结果。分析表明,神经架构在识别上下文无关语言方面的表达能力不受词元化影响。

原文摘要 · Abstract (English)

Tokenization is an important preprocessing step in the training and inference of large language models (LLMs). While there has been extensive research on the expressive power of the neural achitectures used in LLMs, the impact of tokenization has not been well understood. In this work, we demonstrate that tokenization, irrespective of the algorithm used, acts as an inverse homomorphism between strings and tokens. This suggests that the character space of the source language and the token space of the tokenized language are homomorphic, preserving the structural properties of the source language. Additionally, we explore the concept of proper tokenization, which refers to an unambiguous tokenization returned from the tokenizer. Our analysis reveals that the expressiveness of neural architectures in recognizing context-free languages is not affected by tokenization.

词元化形式语言同态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。