arXiv:2506.12928cs.AI2025-06被引 46

通过扩展推理时间计算,提升语言智能体的决策能力

Scaling Test-time Compute for LLM Agents

  • 采用并行采样与串行修正等策略扩展推理时间
  • 增加多样化推演路径可显著提升任务表现
  • 列表式验证合并法效果最佳,反思时机关键

测试时扩展计算已被证明能显著提升大语言模型的推理能力。本文首次系统探索将测试时扩展方法应用于语言智能体,并研究其对智能体有效性的影响程度。具体考察了四种策略:(1)并行采样算法;(2)串行修订策略;(3)验证器与结果融合方法;(4)推演路径多样化策略。通过细致分析与消融实验发现:(1)扩展测试时计算可提升智能体性能;(2)智能体需掌握反思时机;(3)在多种验证与合并方法中,列表式方法表现最优;(4)增加多样化推演路径对任务表现有正向影响。

原文摘要 · Abstract (English)

Scaling test time compute has shown remarkable success in improving the reasoning abilities of large language models (LLMs). In this work, we conduct the first systematic exploration of applying test-time scaling methods to language agents and investigate the extent to which it improves their effectiveness. Specifically, we explore different test-time scaling strategies, including: (1) parallel sampling algorithms; (2) sequential revision strategies; (3) verifiers and merging methods; (4)strategies for diversifying rollouts.We carefully analyze and ablate the impact of different design strategies on applying test-time scaling on language agents, and have follow findings: 1. Scaling test time compute could improve the performance of agents. 2. Knowing when to reflect is important for agents. 3. Among different verification and result merging approaches, the list-wise method performs best. 4. Increasing diversified rollouts exerts a positive effect on the agent's task performance.

语言模型智能体推理优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。