arXiv:2511.09005cs.AIcs.CL2025-11

用历史人物角色模拟,让AI逐步优化推理,效果显著提升。

AI Founding Fathers: A Case Study of GIS Search in Multi-Agent Pipelines

  • 构建多智能体管道,让AI像人类一样逐步迭代改进观点。
  • 复杂模型平均得分88.3,远超简单模型的71.7分。
  • 适合想提升AI逻辑推理能力的研究者与开发者。

尽管大语言模型表现出卓越的语言流畅性,但提升其推理能力仍是关键挑战。本文提出一种系统性框架,认为高质量推理本质上是受控的、渐进的、序列化的搜索过程(GIS)。为验证该框架,我们测试了递归精炼(RR)——一种通过自我批判、对抗性压力测试和反馈整合实现的迭代方法——在多智能体管道中的应用。模型以三位美国开国元勋(汉密尔顿、杰斐逊、麦迪逊)的历史形象为基础,结合RAG增强语料库,回应三个当代政治议题。采用双层评估:由大语言模型仲裁代理打分,辅以人工定性判断。结果表明,复杂结构模型在全部九个案例中均优于简单线性模型,平均得分88.3对比71.7;其论证在分析深度、结构细节和策略布局上更优。结论支持递归精炼作为强化大模型推理能力的有效架构机制。

原文摘要 · Abstract (English)

Although Large Language Models (LLMs) show exceptional fluency, efforts persist to extract stronger reasoning capabilities from them. Drawing on search-based interpretations of LLM computation, this paper advances a systematic framework for understanding LLM reasoning and optimization. Namely, that enhancing reasoning is best achieved by structuring a multi-agent pipeline to ensure a traversal of the search space in a gradual, incremental, and sequential (GIS) manner. Stated succinctly, high-quality reasoning is a controlled, incremental search. To test this framework, we investigate the efficacy of recursive refinement (RR)--an iterative process of self-criticism, adversarial stress-testing, and integrating critical feedback--as a practical method for implementing GIS search. We designed an experiment comparing a simple, linear pipeline against a complex, explicitly structured pipeline leveraging a recursive refinement layer. The multi-agent models were constructed to reflect the historical personas of three US Founding Fathers (Hamilton, Jefferson, and Madison) using RAG-powered corpora and were prompted to generate responses to three contemporary political issues. Model performance was evaluated using a two-tiered approach: a quantitative score from an LLM arbiter agent and qualitative human judgment. Our results revealed that the complex model consistently outperformed the simple model across all nine test cases with an average arbiter-outputted score of 88.3 versus 71.7. The complex model's arguments were superior in analytical depth, structural nuance, and strategic framing. We conclude that recursive refinement is a robust architectural feature for enhancing LLM reasoning via GIS search.

大模型推理多智能体递归精炼

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。