通过筛选高分生成结果,提升大模型推理效率与质量
On the Limits of Test-Time Compute: Sequential Reward Filtering for Better Inference
- 用高奖励生成内容过滤低质输出,聚焦优质候选
- 在多个基准上表现优于主流方法,稳定提升效果
- 理论证明更优,适合追求推理性能的开发者
测试时计算(TTC)已成为增强大语言模型性能的重要范式。尽管最佳- n(BoN)采样和序列修订等方法在实践中取得成功,但其根本局限尚不明确。本文分析了一种参考混合策略模型,证明标准BoN本质上次优。为逼近最优边界,我们研究了奖励过滤的序列推理:仅将高奖励生成内容纳入上下文。该机制将计算集中在优质候选上,抑制劣质输出。理论上,奖励过滤的序列推理比标准TTC范式提供更强保证;实证上,在多个基准上评估该策略,均持续优于广泛使用的现有方法,验证了框架的实用性。
原文摘要 · Abstract (English)
Test-time compute (TTC) has become an increasingly prominent paradigm for enhancing large language models (LLMs). Despite the empirical success of methods such as best-of-$n$ (BoN) sampling and sequential revision, their fundamental limits remain unclear. We address this gap by analyzing a mixture-of-reference policy model and proving that standard BoN is inherently suboptimal. To move closer to the optimal frontier, we study reward-filtered sequential inference, a simple procedure that selectively incorporates only high-reward generations into the context. This mechanism concentrates computation on superior policy candidates and suppresses inferior ones. On the theoretical side, we show that reward-filtered sequential inference yields strictly stronger guarantees than standard TTC paradigms. On the empirical side, we evaluate such an inference strategy across diverse benchmarks and observe consistent improvements over widely used approaches, demonstrating the practical effectiveness of our framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。