arXiv:2506.10268cs.CLcs.AI2025-06被引 3

发现大模型在特定条件下会表现确定性决策,挑战其概率推理假说。

Do Language Models Have Bayesian Brains? Distinguishing Stochastic and Deterministic Decision Patterns within Large Language Models

  • 通过模拟吉布斯采样检测模型是否真在随机采样
  • 高温度下仍可能输出最大似然结果,呈现确定性行为
  • 为避免误推模型先验,提出区分随机与确定决策的方法

语言模型本质上是词元序列的概率分布。自回归模型通过迭代计算并采样下一个词元的分布来生成句子,这种迭代采样引入了随机性,使人们普遍认为语言模型做的是概率决策,类似于从未知分布中采样。受此启发,先前研究采用模拟吉布斯采样方法,借鉴人类先验探测实验,试图推断语言模型的先验。本文重新审视关键问题:语言模型是否有贝叶斯大脑?研究发现,在某些条件下,语言模型可表现出近乎确定性的决策行为,例如输出最大似然估计,即使采样温度非零。这质疑了采样假设,并削弱了此前推断人类类似先验的方法的有效性。此外,我们证明:若缺乏严格检验,一个具有确定性行为的系统在模拟吉布斯采样下也可能收敛到“虚假先验”。为此,我们提出一种简单有效的方法,用于区分吉布斯采样中的随机与确定性决策模式,防止错误推断语言模型先验。我们在多种大型语言模型上进行实验,揭示其在不同情境下的决策模式,为理解大模型决策机制提供关键洞见。

原文摘要 · Abstract (English)

Language models are essentially probability distributions over token sequences. Auto-regressive models generate sentences by iteratively computing and sampling from the distribution of the next token. This iterative sampling introduces stochasticity, leading to the assumption that language models make probabilistic decisions, similar to sampling from unknown distributions. Building on this assumption, prior research has used simulated Gibbs sampling, inspired by experiments designed to elicit human priors, to infer the priors of language models. In this paper, we revisit a critical question: Do language models possess Bayesian brains? Our findings show that under certain conditions, language models can exhibit near-deterministic decision-making, such as producing maximum likelihood estimations, even with a non-zero sampling temperature. This challenges the sampling assumption and undermines previous methods for eliciting human-like priors. Furthermore, we demonstrate that without proper scrutiny, a system with deterministic behavior undergoing simulated Gibbs sampling can converge to a "false prior." To address this, we propose a straightforward approach to distinguish between stochastic and deterministic decision patterns in Gibbs sampling, helping to prevent the inference of misleading language model priors. We experiment on a variety of large language models to identify their decision patterns under various circumstances. Our results provide key insights in understanding decision making of large language models.

大模型决策贝叶斯推理采样机制先验推断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。