arXiv:2507.12399cs.LGstat.ML2025-07被引 10

揭示验证器缺陷如何影响推理时扩展效果,给出理论边界。

ROC-n-reroll: How verifier imperfection affects test-time scaling

  • 用验证器的ROC曲线几何特性刻画推理性能
  • 固定算力下拒绝采样优于最佳N选一,无限算力下两者趋同
  • 低算力表现无法预测高算力表现,警示实验设计

推理时扩展旨在通过增加推理阶段计算量来提升语言模型性能。许多研究已实证探索利用验证器实现推理时扩展的技术,如最佳N选一(BoN)和拒绝采样(RS)。然而,目前对验证器不完美性如何影响性能尚缺乏理论理解。本文填补了这一空白:我们证明,这些方法的实例级准确率可被验证器的ROC曲线几何精确刻画。理论得出两个重要结论,经在GSM8K和MATH500数据集上使用Qwen与LLama模型的实验验证。第一,在固定计算资源下,拒绝采样优于最佳N选一;两者在无限计算极限下收敛至相同准确率。第二,通常无法基于低计算条件下的观测结果预测高计算条件下的性能表现。

原文摘要 · Abstract (English)

Test-time scaling aims to improve language model performance by leveraging additional compute during inference. Many works have empirically studied techniques such as Best-of-N (BoN) and Rejection Sampling (RS) that make use of a verifier to enable test-time scaling. However, to date there is little theoretical understanding of how verifier imperfection affects performance -- a gap we address in this work. Specifically, we prove that the instance-level accuracy of these methods is precisely characterized by the geometry of the verifier's ROC curve. Our theory has two important takeaways, confirmed by experiments with Qwen and LLama models on GSM8K and MATH500. First, RS outperforms BoN for fixed compute, while both methods converge to the same accuracy in the infinite-compute limit. Second, it is generally impossible to predict the high-compute performance of either method based on observations in the low-compute regime.

推理优化验证器理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。