让大模型零样本推理更准:通过结构分析引导思维过程
Make LLMs better zero-shot reasoners: Structure-orientated autonomous reasoning
- 用结构化分析解析问题,引导大模型分步推理
- 在复杂任务上超越现有零样本方法,部分场景接近少样本效果
- 多智能体系统支持纠错和查知识,抗干扰能力强
零样本推理方法具有强大的泛化能力且无需人工示例,但在多步推理等复杂任务中仍存在不足。本文提出一种面向结构的分析方法,帮助大模型更好理解问题并引导求解过程。我们验证了该方法对现有链式思考与ReAct策略的提升作用,并借助概率图模型从理论上解释其有效性。为进一步提高可靠性,提出多智能体系统SARA,通过精炼技术和外部知识检索减少事实错误。大量实验表明,该系统不仅在复杂问答中显著提升准确率,还对干扰推理过程的攻击具备鲁棒性,某些情况下甚至超越少样本方法。
原文摘要 · Abstract (English)
Zero-shot reasoning methods with Large Language Models (LLMs) offer significant advantages including great generalization to novel tasks and reduced dependency on human-crafted examples. However, the current zero-shot methods still have limitations in complex tasks, e.g., answering questions that require multi-step reasoning. In this paper, we address this limitation by introducing a novel structure-oriented analysis method to help LLMs better understand the question and guide the problem-solving process of LLMs. We first demonstrate how the existing reasoning strategies, Chain-of-Thought and ReAct, can benefit from our structure-oriented analysis. In addition to empirical investigations, we leverage the probabilistic graphical model to theoretically explain why our structure-oriented analysis can improve the LLM reasoning process. To further improve the reliability in complex question-answering tasks, we propose a multi-agent reasoning system, Structure-oriented Autonomous Reasoning Agents (SARA), that can better enforce the reasoning process following our structure-oriented analysis by refinement techniques and is equipped with external knowledge retrieval capability to reduce factual errors. Extensive experiments verify the effectiveness of the proposed reasoning system. Surprisingly, in some cases, the system even surpasses few-shot methods. Finally, the system not only improves reasoning accuracy in complex tasks but also demonstrates robustness against potential attacks that corrupt the reasoning process.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。