arXiv:2608.14665cs.LG2026-08

解释为何采样预算越大,最优温度越高。

When Does the Best Sampling Temperature Rise with the Budget? Sufficient Conditions for Pass@k

  • 提出一个理论条件:成功概率越低的任务更受高温度青睐。
  • 发现最优温度随预算增加而上升,且在不同任务中表现一致。
  • 适合研究大模型推理策略与温度调优的学者参考。

当采样预算较小时,最大化 pass@k 的最优温度通常较低;预算增大时,该温度则升高。这一现象在 Codex 到最近的多样本推理研究中均有观察。该模式并非 pass@k 的代数特性——如 Slocum 等人(ICLR 2025)指出,对固定任务而言,最优温度与 k 无关。本文基于此固定任务观察与难易任务解释,给出一个群体层面的充分条件。对于任务 X,设 $p_t(X)$ 为温度 t 下的一次采样成功率,定义条件对数成功响应 $m_t(u) = \mathbb{E}[\dot p_t(X) \mid p_t(X)=u]/u$。若 $m_t(u)$ 随当前成功率非增,则归一化温度导数的 aggregate pass@k 随 k 非减。因此,导数符号在不同预算间嵌套;若每条温度-性能曲线严格单峰,则其唯一最大值点随 k 非减。证明揭示机制为单调似然比幂倾斜,偏向低成功率任务。我们推导出闭式两层相图,包含上行与下行区域,并表明边际温度导数具有精确的 $\mathrm{Beta}(2,k)$ 核表示,其核集中在约 $1/k$ 的单样本成功率尺度。将该尺度解释为任务级局域化,还需在零附近存在正则且非消失的密度响应因子。符号矩表示提供诊断形状约束,附录还记录了现有多配置分配公式的精确离散改进。未训练任何语言模型,也无模型查询作为实验测量:贡献在于对已知经验现象的条件性理论,假设可在未来工作中验证。

原文摘要 · Abstract (English)

The temperature that maximizes pass@$k$ is often low for a small sampling budget and higher for a large budget. This pattern has been reported from Codex through recent multi-sample inference studies. It is not an algebraic property of pass@$k$: as Slocum et al. (ICLR 2025) observe, for one fixed task the maximizing temperature is independent of $k$. Building on that fixed-task observation and the hard/easy-task explanation, we give a formal population-level sufficient condition for the aggregate pattern. For task $X$, let $p_t(X)$ be one-sample success probability at temperature $t$, and define the conditional log-success response $m_t(u)=\mathbb{E}[\dot p_t(X)\mid p_t(X)=u]/u$. If $m_t(u)$ is nonincreasing in current success probability, then the normalized temperature derivative of aggregate pass@$k$ is nondecreasing in $k$. Consequently, derivative signs are nested across budgets; if each temperature-performance curve is strictly single-peaked, its unique maximizer is nondecreasing in $k$. The proof identifies the mechanism as a monotone-likelihood-ratio power tilt toward lower-success tasks. We derive a closed-form two-stratum phase diagram, including upward and downward regimes, and show that the marginal temperature derivative admits an exact $\mathrm{Beta}(2,k)$ kernel representation whose kernel concentrates at one-sample success of order $1/k$. Interpreting that scale as task-level localization additionally requires a regular, nonvanishing density-response factor near zero. A signed-moment representation yields diagnostic shape restrictions, while a short appendix records exact discrete refinements of the existing multi-configuration allocation formulation. No language model is trained, and no model query is used as an experimental measurement: the contribution is a conditional theory of an established empirical phenomenon, with assumptions that can be tested in future work.

温度调优推理策略概率分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。