提出背景温度概念,揭示大模型推理时的隐性随机性来源。
Introducing Background Temperature to Characterise Hidden Randomness in Large Language Models

- 引入背景温度 $T_{\mathrm{bg}}$ 揭示实现层面的隐性扰动
- 实验证明即使 $T=0$ 仍存在输出差异,$T_{\mathrm{bg}}$ 可量化
- 适用于关注模型可复现性与部署稳定性的研究者
即便在温度 $T=0$ 的情况下,大型语言模型(LLMs)对相同输入仍可能产生不同输出。近期思维机器实验室的研究指出,非确定性来源于实现层面,包括批大小变化、核函数非不变性以及浮点数非结合性。本文通过引入‘背景温度’$T_{\mathrm{bg}}$ 概念,形式化描述了当名义温度为零时,由实现依赖扰动所导致的有效温度。我们给出了清晰定义,阐明 $T_{\mathrm{bg}}$ 与推理环境 $I$ 所决定的随机扰动之间的关系,并提出一种基于理想参考系统等效温度 $T_n(I)$ 的经验估计方法。最后,在主流 LLM 提供商的代表性模型上进行了初步实验,验证了该方法的有效性,并讨论了其对可复现性、评估和部署的影响。
原文摘要 · Abstract (English)
Even when decoding with temperature $T=0$, large language models (LLMs) can produce divergent outputs for identical inputs. Recent work by Thinking Machines Lab highlights implementation-level sources of nondeterminism, including batch-size variation, kernel non-invariance, and floating-point non-associativity. In this short note we formalize this behavior by introducing the notion of \emph{background temperature} $T_{\mathrm{bg}}$, the effective temperature induced by an implementation-dependent perturbation process observed even when nominal $T=0$. We provide clean definitions, show how $T_{\mathrm{bg}}$ relates to a stochastic perturbation governed by the inference environment $I$, and propose an empirical protocol to estimate $T_{bg}$ via the equivalent temperature $T_n(I)$ of an ideal reference system. We conclude with a set of pilot experiments run on a representative pool from the major LLM providers that demonstrate the idea and outline implications for reproducibility, evaluation, and deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。