用多个专家+批评者协同纠错,提升数学推理可靠性
Critic-Guided Heterogeneous Multi-Agent Reasoning for Reliable Mathematical Problem Solving
- 多角色协作:不同专长的LLM分工,由批评者动态评估并指导修正
- 在GSM8K上比单次推理模型高13%准确率,小模型也能达大模型水平
- 批评反馈环是关键,不依赖模型大小,适合追求可靠性的实际应用
近期大型语言模型在复杂数学推理中展现出强大能力,但仍易出现幻觉、中间推理错误及不可靠结果。本文提出一种基于批评者的异构多智能体框架,通过多个具有不同专长的LLM智能体协作,并引入批评者驱动的自适应学习机制,根据中间反馈评估并引导推理过程。系统采用生成-验证架构,验证者不仅判断答案正确性,还提供批判性反馈以指导解题重生成,实现自适应纠错并防止错误传播。在GSM8K基准测试中,该方法相比单次推理和非批评模型最高提升13%准确率。消融实验表明,性能提升主要源于批评反馈循环,而非模型规模。研究显示,异构协作与批评机制可显著降低对大模型的依赖,使小模型表现达到媲美水平。整体表明,融合异构多智能体与批判反馈能构建更可靠、可解释的推理系统。
原文摘要 · Abstract (English)
Recent Large Language Models (LLMs) have shown impressive reasoning abilities; but they are still susceptible to hallucinations, intermediate reasoning mistakes, and unreliable reasoning results in complex mathematical reasoning problems. In this study, we introduce a critic-based heterogeneous multi-agent approach to improve the dependability of mathematical reasoning. This framework incorporates several LLM agents of different specialties and employs a critic-driven adaptive learning system to assess and guide the reasoning process based on intermediate feedback. The system adopts a generator-validator framework, with the validator not only determining correctness but also offering critiques to guide regeneration of solutions. This allows for adaptive error correction and prevents error cascading. Our experiments on the GSM8K benchmark show that the proposed method achieves up to 13% accuracy improvement over single-shot and non-critic models. Additionally, findings suggest that heterogeneity and critique reduce the need for large models, allowing smaller models to perform on par. Ablation studies reveal the main performance gains are due to the critic-based feedback loop and not model size. In summary, the proposed approach showcases the benefits of combining heterogeneous multi-agent collaboration and critique to obtain reliable and interpretable reasoning systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。