arXiv:2506.08507cs.MAcs.AI2025-06被引 2

用强化学习自动构建智能体系统,让多个智能体自主协作完成任务。

MasHost Builds It All: Autonomous Multi-Agent System Directed by Reinforcement Learning

  • 将智能体系统构建视为图搜索问题,用统一概率采样生成角色与互动关系。
  • 在6个基准上表现优于主流方法,准确率和效率均显著提升。
  • 首次实现全自主智能体系统设计,适合复杂任务自动化研究者。

基于大语言模型的多智能体系统(Mas)已成为解决复杂现实任务的强大范式。然而,现有构建方法通常依赖人工设计的交互机制或启发式规则,引入人为偏见并限制自主性。即使近期出现自适应构建进展,多数系统仍处于半自主模式。本文提出MasHost,一种基于强化学习(RL)的自主、查询自适应多智能体系统设计框架。通过将系统构建建模为图搜索问题,MasHost采用统一的概率采样机制联合生成智能体角色及其交互关系。除精度与效率外,我们引入组件合理性作为新的设计原则。为实现多目标优化,提出分层相对策略优化(HRPO),协同整合群体相对优势与动作级奖励。据我们所知,MasHost是首个基于强化学习的自主多智能体图结构构建框架。在六个基准上的大量实验表明,其持续优于多数竞争基线,验证了其有效性、高效性与结构合理性。

原文摘要 · Abstract (English)

Large Language Model (LLM)-driven Multi-agent systems (Mas) have recently emerged as a powerful paradigm for tackling complex real-world tasks. However, existing Mas construction methods typically rely on manually crafted interaction mechanisms or heuristic rules, introducing human biases and constraining the autonomous ability. Even with recent advances in adaptive Mas construction, existing systems largely remain within the paradigm of semi-autonomous patterns. In this work, we propose MasHost, a Reinforcement Learning (RL)-based framework for autonomous and query-adaptive Mas design. By formulating Mas construction as a graph search problem, our proposed MasHost jointly samples agent roles and their interactions through a unified probabilistic sampling mechanism. Beyond the accuracy and efficiency objectives pursued in prior works, we introduce component rationality as an additional and novel design principle in Mas. To achieve this multi-objective optimization, we propose Hierarchical Relative Policy Optimization (HRPO), a novel RL strategy that collaboratively integrates group-relative advantages and action-wise rewards. To our knowledge, our proposed MasHost is the first RL-driven framework for autonomous Mas graph construction. Extensive experiments on six benchmarks demonstrate that MasHost consistently outperforms most competitive baselines, validating its effectiveness, efficiency, and structure rationality.

多智能体强化学习自主系统大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。