arXiv:2504.14520cs.AIcs.CL2025-04综述被引 12

用多智能体强化学习让大模型学会自我反思,提升可靠性。

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey

  • 设计多智能体架构模拟人类内省,如监督层级与智能体辩论
  • 通过自洽反馈和持续学习机制增强模型自我评估能力
  • 适合研究可信AI、复杂决策系统构建的学者与工程师

本文从多智能体强化学习(MARL)视角综述大语言模型(LLM)元思维能力的发展。当前LLM存在幻觉频发、缺乏内部自评估机制等问题。现有方法如基于人类反馈的强化学习(RLHF)、自蒸馏和思维链提示均有局限。本文重点探讨多智能体架构——包括监督者-执行者层级、智能体辩论及心智理论框架——如何模拟人类内省行为,提升模型鲁棒性。通过分析奖励机制、自对弈与持续学习方法,提出构建具有自我反思、自适应与可信赖能力的LLM的综合路径。同时讨论了评估指标、数据集及未来方向,如神经科学启发架构与符号推理融合。

原文摘要 · Abstract (English)

This survey explores the development of meta-thinking capabilities in Large Language Models (LLMs) from a Multi-Agent Reinforcement Learning (MARL) perspective. Meta-thinking self-reflection, assessment, and control of thinking processes is an important next step in enhancing LLM reliability, flexibility, and performance, particularly for complex or high-stakes tasks. The survey begins by analyzing current LLM limitations, such as hallucinations and the lack of internal self-assessment mechanisms. It then talks about newer methods, including RL from human feedback (RLHF), self-distillation, and chain-of-thought prompting, and each of their limitations. The crux of the survey is to talk about how multi-agent architectures, namely supervisor-agent hierarchies, agent debates, and theory of mind frameworks, can emulate human-like introspective behavior and enhance LLM robustness. By exploring reward mechanisms, self-play, and continuous learning methods in MARL, this survey gives a comprehensive roadmap to building introspective, adaptive, and trustworthy LLMs. Evaluation metrics, datasets, and future research avenues, including neuroscience-inspired architectures and hybrid symbolic reasoning, are also discussed.

元思维多智能体大模型自反思

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。