梳理数学推理中分步与结果反馈提升大模型能力的方法
A Survey on Feedback-based Multi-step Reasoning for Large Language Models on Mathematics
- 用步骤或结果奖励让大模型自我修正推理过程
- 冻结模型+外部工具可低成本实现反馈增强
- 适合研究大模型数学推理与自洽性改进的学者
大语言模型在数学推理方面取得进展,链式思维提示策略通过多步推理提升解决问题能力。后续研究尝试将多步推理过程融入模型内部,通过过程奖励作为反馈实现优于提示策略的效果。由于步骤级标注成本高,部分工作转向以结果奖励为反馈。此外,无需训练的技术利用冻结的LLM或外部工具,在每一步提供反馈以增强推理。鉴于数学问题逻辑性强,相关研究丰富,本文系统综述了在步骤和结果层面利用反馈提升大模型多步数学推理能力的策略。随着多步推理成为大模型扩展的关键组件,本综述旨在建立理解基础,推动进一步研究。
原文摘要 · Abstract (English)
Recent progress in large language models (LLM) found chain-of-thought prompting strategies to improve the reasoning ability of LLMs by encouraging problem solving through multiple steps. Therefore, subsequent research aimed to integrate the multi-step reasoning process into the LLM itself through process rewards as feedback and achieved improvements over prompting strategies. Due to the cost of step-level annotation, some turn to outcome rewards as feedback. Aside from these training-based approaches, training-free techniques leverage frozen LLMs or external tools for feedback at each step to enhance the reasoning process. With the abundance of work in mathematics due to its logical nature, we present a survey of strategies utilizing feedback at the step and outcome levels to enhance multi-step math reasoning for LLMs. As multi-step reasoning emerges a crucial component in scaling LLMs, we hope to establish its foundation for easier understanding and empower further research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。