arXiv:2410.01920cs.LG2024-10ICLR被引 20

用扭曲序贯蒙特卡洛提升大模型解数学题的效率和准确性

Step-by-Step Reasoning for Math Problems via Twisted Sequential Monte Carlo

  • 通过扭曲序贯蒙特卡洛动态聚焦优质解题路径
  • 在多个数学基准上显著减少采样次数并提高正确率
  • 无需逐步人工标注,适合需要高效推理的场景

增强大语言模型(LLMs)的多步推理能力仍是持续挑战。近期验证机制通过评估生成结果提升了答案一致性,但现有方法存在采样效率低的问题,需大量样本才能达到理想性能。此外,训练有效验证器通常依赖昂贵的全过程监督。本文提出基于扭曲序贯蒙特卡洛(TSMC)的新验证方法,通过在部分解的基础上估计未来期望奖励,逐步优化采样策略,集中探索有潜力的候选解。该方法简化了训练目标,无需逐步人工标注。我们在多个数学基准上验证了其有效性,并支持了理论分析,证明了本方法及现有方法的优势。

原文摘要 · Abstract (English)

Augmenting the multi-step reasoning abilities of Large Language Models (LLMs) has been a persistent challenge. Recently, verification has shown promise in improving solution consistency by evaluating generated outputs. However, current verification approaches suffer from sampling inefficiencies, requiring a large number of samples to achieve satisfactory performance. Additionally, training an effective verifier often depends on extensive process supervision, which is costly to acquire. In this paper, we address these limitations by introducing a novel verification method based on Twisted Sequential Monte Carlo (TSMC). TSMC sequentially refines its sampling effort to focus exploration on promising candidates, resulting in more efficient generation of high-quality solutions. We apply TSMC to LLMs by estimating the expected future rewards at partial solutions. This approach results in a more straightforward training target that eliminates the need for step-wise human annotations. We empirically demonstrate the advantages of our method across multiple math benchmarks, and also validate our theoretical analysis of both our approach and existing verification methods.

大模型推理数学问题蒙特卡洛验证机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。