arXiv:2508.10018cs.CLcs.AI2025-08被引 1

用范畴同伦理论让大模型对同义句生成相同概率输出

A Rose by Any Other Name Would Smell as Sweet: Categorical Homotopy Theory for Large Language Models

  • 构建语言生成的范畴马尔可夫模型,用箭头表示句子概率
  • 引入同伦等价概念解决同义句产生不同输出的问题
  • 理论性强,适合研究模型语义一致性与形式化推理的学者

自然语言中存在大量表面不同但语义相同的表达,如“查尔斯·达尔文写了”和“查尔斯·达尔文是作者”。大语言模型(LLMs)本应对此类句子生成相同的下一个词概率,但实际常出现差异。已有工作尝试通过k-NN相似度估计进行平滑处理。本文从更抽象的角度切入,提出一种用于大语言模型的范畴同伦框架。我们引入一个LLM马尔可夫范畴来表示语言生成中的概率分布,其中句子“查尔斯·达尔文写了”的概率由范畴中的一个箭头定义。然而,由于语言中存在大量等价改写,每种表达都会生成非同构的箭头,导致问题复杂化。为此,我们采用范畴同伦技术,捕捉该范畴内的“弱等价”关系。本文系统阐述了范畴同伦在大语言模型中的应用,涵盖高阶代数K理论、模型范畴等近半个世纪发展的强大理论成果。

原文摘要 · Abstract (English)

Natural language is replete with superficially different statements, such as ``Charles Darwin wrote" and ``Charles Darwin is the author of", which carry the same meaning. Large language models (LLMs) should generate the same next-token probabilities in such cases, but usually do not. Empirical workarounds have been explored, such as using k-NN estimates of sentence similarity to produce smoothed estimates. In this paper, we tackle this problem more abstractly, introducing a categorical homotopy framework for LLMs. We introduce an LLM Markov category to represent probability distributions in language generated by an LLM, where the probability of a sentence, such as ``Charles Darwin wrote" is defined by an arrow in a Markov category. However, this approach runs into difficulties as language is full of equivalent rephrases, and each generates a non-isomorphic arrow in the LLM Markov category. To address this fundamental problem, we use categorical homotopy techniques to capture ``weak equivalences" in an LLM Markov category. We present a detailed overview of application of categorical homotopy to LLMs, from higher algebraic K-theory to model categories, building on powerful theoretical results developed over the past half a century.

大模型范畴论同伦理论语义一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。