arXiv:2503.05188cs.CLcs.AI2025-03被引 3

改进推理时奖励模型,让大模型答得更准更稳

Fixing the Broken Compass: Diagnosing and Improving Inference-Time Reward Modeling

  • 按最终答案聚类推理路径,分组整合奖励信号
  • 实验显示准确率最高提升5%,平均比先进模型高10%
  • 适合想提升大模型推理能力的研究者和工程师

推理时扩展技术在增强大语言模型(LLM)推理能力方面展现出潜力。尽管现有研究多关注训练时优化,本文强调推理时基于奖励模型(RM)的推理是一个关键但被忽视的方向。我们系统分析了下游推理任务中RM的行为,发现三大局限:(1) RM可能损害简单问题的表现;(2) 随着采样数量增加,其判别能力下降;(3) 高搜索多样性会削弱RM性能。为此,我们提出CRISP(分步前缀聚合的聚类奖励集成)算法,通过按最终答案对生成路径聚类,在簇层面聚合奖励信号,并自适应更新前缀提示以引导生成。实验表明,CRISP显著提升了LLM推理性能,相比其他基于RM的推理方法最高提升5%准确率,平均比先进推理模型高出10%。

原文摘要 · Abstract (English)

Inference-time scaling techniques have shown promise in enhancing the reasoning capabilities of large language models (LLMs). While recent research has primarily focused on training-time optimization, our work highlights inference-time reward model (RM)-based reasoning as a critical yet overlooked avenue. In this paper, we conduct a systematic analysis of RM behavior across downstream reasoning tasks, revealing three key limitations: (1) RM can impair performance on simple questions, (2) its discriminative ability declines with increased sampling, and (3) high search diversity undermines RM performance. To address these issues, we propose CRISP (Clustered Reward Integration with Stepwise Prefixing), a novel inference-time algorithm that clusters generated reasoning paths by final answers, aggregates reward signals at the cluster level, and adaptively updates prefix prompts to guide generation. Experimental results demonstrate that CRISP significantly enhances LLM reasoning performance, achieving up to 5% accuracy improvement over other RM-based inference methods and an average of 10% gain over advanced reasoning models.

推理增强奖励模型大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。