用搜索和自反馈提升智能体推理能力,发现真实反馈更有效。
A Study on Leveraging Search and Self-Feedback for Agent Reasoning
- 结合搜索与模型自反馈优化推理路径探索
- 自反馈在数学推理中效果有限,泛化能力差
- 复杂任务需设计专用反馈机制,或依赖真实反馈
近期研究显示,推理阶段引入搜索可显著提升语言智能体的推理能力。部分方法依赖真实反馈或模型自生成反馈,通过反馈调整搜索策略以优化探索与利用路径。本研究探讨搜索与模型自反馈在推理任务中的协同作用。首先,对比了数学推理中真实反馈与自反馈的差异;其次,发现搜索技术在工具调用和设计等复杂任务中存在局限性,并提出领域特定的改进方案。实验表明,仅依赖自反馈时存在泛化挑战,为使搜索有效,需有真实反馈支持,或针对具体任务精心设计反馈机制。
原文摘要 · Abstract (English)
Recent works have demonstrated that incorporating search during inference can significantly improve reasoning capabilities of language agents. Some approaches may make use of the ground truth or rely on model's own generated feedback. The search algorithm uses this feedback to then produce values that will update its criterion for exploring and exploiting various reasoning paths. In this study, we investigate how search and model's self-feedback can be leveraged for reasoning tasks. First, we explore differences in ground-truth feedback and self-feedback during search for math reasoning. Second, we observe limitations in applying search techniques to more complex tasks like tool-calling and design domain-specific approaches to address these gaps. Our experiments reveal challenges related to generalization when solely relying on self-feedback during search. For search to work effectively, either access to the ground-truth is needed or feedback mechanisms need to be carefully designed for the specific task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。