通过两次调用揭示大模型推理的稳定性,精准预测多次投票效果。
Two Calls, Two Moments, and the Vote-Accuracy Curve of Repeated LLM Inference

- 利用两次调用估计正确率均值与方差,区分固定错误与随机可恢复误差。
- 三票投票准确率区间宽度不超过1/8,五票结果在理论范围内。
- 适用于评估模型鲁棒性,尤其适合对比不同温度或混合策略的效果。
重复采样是测试时计算资源的标准使用方式,但其收益受样本间正确性潜在分布制约,而非仅由单次调用准确率决定。我们在条件独立同分布假设下研究二值正确性层的重复大模型推理。一次标注调用可识别平均潜在成功概率;两次标注调用可识别其二阶矩,从而确定相同样本的正确性相关性,以区分稳定错误与可恢复的调用级随机性。基于这两个矩,每个固定的多数投票预算均有严格无分布的双调用区间。关键技术还原为:无穷维矩问题具有三原子极值解,且对任意有限预算存在二次对偶证书,因此边界为精确值而非离散化或参数化近似。首个实用预算(三票)有闭式解,区间宽度不超过1/8,并具备可验证改进标准。无限投票极限即多数投票随调用数趋近无穷时的极限,虽仍被严格界定,但对潜质在q=1/2附近的质量敏感。我们引入最大熵与潜难度高斯-概率点补全方法,实验在QNLI和QQP上的大模型调用表明,经验三票与五票准确率均包含于投影后的双调用区域中;温度变化与随机模型混合可产生不依赖单次准确率的投票增益。
原文摘要 · Abstract (English)
Repeated sampling is a standard way to spend test-time compute, but its benefit is controlled by the latent distribution of correctness across examples, not by one-call accuracy alone. We study the binary correctness layer of repeated LLM inference under conditional-i.i.d. calls. One labeled call identifies the mean latent success probability; two labeled calls identify its second moment and hence the same-example correctness correlation that separates stable errors from recoverable call-level randomness. From these two moments, every fixed majority-vote budget has a sharp distribution-free two-call interval. The key technical reduction is that the infinite-dimensional moment problem has three-atom extremizers and quadratic dual certificates for every finite budget, so the bounds are exact rather than discretized or parametric. The first useful budget, three votes, has a closed form, width at most $1/8$, and a certified-improvement criterion. The infinite-vote endpoint is the limit of majority voting as the number of calls tends to infinity; it is also sharply bounded, but remains threshold-sensitive because it depends on latent mass around $q=1/2$. We add maximum-entropy and Latent-difficulty Gaussian-probit point completions, and experiments on LLM calls over QNLI and QQP show that empirical three- and five-vote accuracies are contained in the projected two-call regions while temperature changes and randomized model mixtures can create voting gains not ordered by one-call accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。