arXiv:2603.07972cs.AI2026-03被引 7

让AI团队学会何时自决、何时求助人类,持续进化协作能力。

Adaptive Collaboration with Humans: Metacognitive Policy Optimization for Multi-Agent LLMs with Continual Learning

  • 设计元认知策略,判断何时自主解决、何时请求人类帮助。
  • 双环优化机制使系统在任务中决策,并通过人类反馈持续提升能力。
  • 在数学与推理任务上超越现有AI团队,适合需要长期进化的协作场景。

尽管单个大语言模型的规模扩展带来了显著进展,但下一代突破在于通过多智能体系统(MAS)实现协作规模化。然而,完全自主的MAS仍属于‘封闭世界’系统,受限于预训练模型的静态知识范围,在面对训练数据之外的知识时极易失效,导致集体失败。为此,我们提出‘人在回路中的多智能体协作’(HILA)框架,一种人机协同的原理性范式。HILA训练智能体学习元认知策略,以决定何时自主解决问题,何时向人类专家求助。为实现该策略,我们引入双环策略优化:内环采用带成本感知奖励的组相对策略优化(GRPO),优化求助决策;外环实施持续学习,将专家反馈转化为高质量监督信号,增强智能体推理能力。在具有挑战性的数学与问题求解基准测试中,配备双环优化的HILA持续优于先进多智能体系统,为可协作、可持续进化的智能体系统奠定了原理基础。

原文摘要 · Abstract (English)

While scaling individual Large Language Models (LLMs) has delivered remarkable progress, the next frontier lies in scaling collaboration through multi-agent systems (MAS). However, purely autonomous MAS remain ''closed-world'' systems, constrained by the static knowledge horizon of pre-trained models. This limitation makes them brittle on tasks requiring knowledge beyond training data, often leading to collective failure under novel challenges. To address this, we propose the Human-In-the-Loop Multi-Agent Collaboration (HILA) framework, a principled paradigm for human--agent collaboration. HILA trains agents to learn a metacognitive policy that governs when to solve problems autonomously and when to defer to a human expert. To operationalize this policy, we introduce Dual-Loop Policy Optimization, which disentangles immediate decision-making from long-term capability growth. The inner loop applies Group Relative Policy Optimization (GRPO) with a cost-aware reward to optimize deferral decisions, while the outer loop implements continual learning, transforming expert feedback into high-quality supervised signals that strengthen the agent's reasoning ability. Experiments on challenging mathematical and problem-solving benchmarks show that HILA, equipped with Dual-Loop Policy Optimization, consistently outperforms advanced MAS, establishing a principled foundation for collaborative and continually improving agentic systems.

多智能体人机协作持续学习元认知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。