arXiv:2506.00396cs.CL2025-06ACL被引 1

用低成本提升大模型决策能力,效率提高10倍

Speculative Reward Model Boosts Decision Making Ability of LLMs Cost-Effectively

  • 引入外部奖励模型预测最优动作,减少大模型自评依赖
  • 通过推测验证机制剪枝劣解,平均降低90%计算成本
  • 适合需要高效推理的数学与规划类任务场景

大语言模型的有效决策对处理复杂任务至关重要。现有方法虽注重性能,却常忽视效果与计算成本的平衡。为此,我们提出3E标准系统评估搜索策略的成本效益,发现现有方法常以显著效率损失换取微弱性能提升。为在保持效率的同时增强决策能力,我们提出可插拔的推测奖励模型(SRM)框架,该框架通过外部奖励分配器预测最优动作,减少对大模型内部自评价的依赖;并引入推测验证机制,剪枝次优选择,引导搜索向更有希望的步骤推进。我们在数学推理、规划及专业领域数值推理等复杂任务上评估了SRM,实验表明其平均计算成本降至原框架的1/10,同时保持原有有效性。

原文摘要 · Abstract (English)

Effective decision-making in Large Language Models (LLMs) is essential for handling intricate tasks. However, existing approaches prioritize performance but often overlook the balance between effectiveness and computational cost. To address this, we first introduce the 3E Criteria to systematically assess the cost-effectiveness of search strategies, revealing that existing methods often trade significant efficiency for marginal performance gains. To improve LLM decision-making while maintaining efficiency, we propose the Speculative Reward Model (SRM), a plug-and-play framework that seamlessly integrates with existing search strategies. Specifically, SRM employs an external reward assigner to predict optimal actions, reducing reliance on LLMs' internal self-evaluation. And a speculative verification mechanism is used to prune suboptimal choices and guide the search toward more promising steps. We evaluate SRM on several complex decision-making tasks including mathematical reasoning, planning and numerical reasoning in specialized domains. Experimental results show that SRM reduces costs to 1/10 of the original search framework on average while maintaining effectiveness.

大模型决策奖励模型高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。