arXiv:2510.13918cs.CL2025-10被引 2

通过智能加权融合大模型与评分模型信号,提升测试时扩展效率。

Optimal Aggregation of LLM and PRM Signals for Efficient Test-Time Scaling

  • 提出理论框架,实现大模型与评分模型信号的最优加权融合。
  • 实验证明最优权重常含负值,且仅需21.3%计算量即超越传统方法。
  • 适合关注测试时扩展效率与模型集成的科研人员与工程师。

过程评分模型(PRM)是测试时扩展(TTS)的核心,用于验证并筛选大语言模型(LLM)的最佳响应。然而,近期基准测试显示,忽略PRM信号的简单多数投票偶尔优于标准的基于PRM的选择方法。这引发关键问题:如何有效利用PRM的验证信号?为此,我们构建了理论上最优的LLM与PRM信号融合框架,揭示最优策略为带权重的响应聚合,其有效性取决于对复杂模型交互的权重估计。基于此,我们实证发现最优权重函数在不同LLM-PRM组合间差异显著,且常赋予负权重。受此启发,我们提出高效预计算方法以校准这些权重。在5个LLM和7个PRM上的广泛实验表明,该校准方法显著提升TTS效率,性能超越原始加权多数投票,同时仅消耗21.3%的计算资源。研究证明,投入更智能的聚合策略比单纯增加测试时计算量更能带来性能提升。

原文摘要 · Abstract (English)

Process reward models (PRMs) are a cornerstone of test-time scaling (TTS), designed to verify and select the best responses from large language models (LLMs). However, this promise is challenged by recent benchmarks where simple majority voting, which ignores PRM signals, occasionally outperforms standard PRM-based selection. This raises a critical question: How can we effectively utilize verification signals from PRMs for TTS? To address this, we start by developing a theoretical framework for optimally combining signals from both the LLM and the PRM. Our framework reveals that the optimal strategy is a weighted aggregation of responses, a strategy whose effectiveness hinges on estimating weights that capture the complex interplay between the models. Based on our theoretical results, we empirically show that these optimal weighting functions differ significantly across LLM-PRM pairs and, notably, often assign substantial negative weights. Motivated by these insights, we propose efficient pre-computation methods to calibrate these weighting functions. Extensive experiments across 5 LLMs and 7 PRMs demonstrate that our calibration method significantly boosts the TTS efficiency, surpassing the performance of vanilla weighted majority voting while using only $21.3\%$ of the computation. Ultimately, our work demonstrates that investing in a more intelligent aggregation strategy can be a more convincing path to performance gains than simply scaling test-time computation.

测试时扩展模型融合评分模型效率优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。