arXiv:2501.06662math.CTcs.CL2025-01被引 7

用语言模型概率定义文本类别,计算其信息量与结构特征。

The magnitude of categories of texts enriched by language models

  • 基于模型预测概率构建文本的数学空间,定义类别丰富度。
  • 揭示文本空间的大小函数与熵、输出数量的关系。
  • 适合对语言模型数学结构感兴趣的理论研究者。

本文从两个方面展开:首先,利用语言模型的下一个词概率,明确地在单位区间内定义自然语言文本的类别丰富度,考虑文本生成的终止条件,并确定该丰富度可解释为文本上的概率分布。其次,计算相关广义度量空间的莫比乌斯函数和大小(magnitude)。该空间的大小函数是所有提示(prompt)的 $t$-对数(Tsallis)熵之和,加上模型可能输出的基数。对大小函数导数的合理评估可恢复香农熵之和,支持将大小视为配分函数。此外,依据莱因斯特与舒尔曼的工作,将大小函数表达为大小同调的欧拉示性数,并给出零阶与一阶大小同调群的显式描述。

原文摘要 · Abstract (English)

The purpose of this article is twofold. Firstly, we use the next-token probabilities given by a language model to explicitly define a category of texts in natural language enriched over the unit interval, in the sense of Bradley, Terilla, and Vlassopoulos. We consider explicitly the terminating conditions for text generation and determine when the enrichment itself can be interpreted as a probability over texts. Secondly, we compute the Möbius function and the magnitude of an associated generalized metric space of texts. The magnitude function of that space is a sum over texts (prompts) of the $t$-logarithmic (Tsallis) entropies of the next-token probability distributions associated with each prompt, plus the cardinality of the model's possible outputs. A suitable evaluation of the magnitude function's derivative recovers a sum of Shannon entropies, which justifies seeing magnitude as a partition function. Following Leinster and Shulman, we also express the magnitude function of the generalized metric space as an Euler characteristic of magnitude homology and provide an explicit description of the zeroeth and first magnitude homology groups.

语言模型信息论数学结构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。