arXiv:2503.13553cs.MAcs.AI2025-03被引 4

用大模型指导多智能体强化学习,提升训练效率与表现

LLM-Mediated Guidance of MARL Systems

  • 用大模型生成自然语言或规则指令干预多智能体学习
  • 早期干预使智能体训练更高效,性能显著优于无干预基线
  • 基于规则的干预比自然语言干预效果更强,适合复杂场景

在复杂的多智能体环境中,实现高效学习和理想行为对多智能体强化学习(MARL)系统构成重大挑战。本文探索将多智能体强化学习与大语言模型(LLM)引导的干预相结合,以引导智能体向更优行为演进。我们研究了两类干预控制器:自然语言(NL)控制器和基于规则(RB)控制器。实验表明,基于规则的控制器比使用小型(7B/8B)LLM模拟人类干预的自然语言控制器影响更显著。研究发现,早期干预能显著提升训练效率与最终性能。两类干预均优于无干预基线,证明了大语言模型引导在加速训练和提升复杂环境下的MARL性能方面具有潜力。

原文摘要 · Abstract (English)

In complex multi-agent environments, achieving efficient learning and desirable behaviours is a significant challenge for Multi-Agent Reinforcement Learning (MARL) systems. This work explores the potential of combining MARL with Large Language Model (LLM)-mediated interventions to guide agents toward more desirable behaviours. Specifically, we investigate how LLMs can be used to interpret and facilitate interventions that shape the learning trajectories of multiple agents. We experimented with two types of interventions, referred to as controllers: a Natural Language (NL) Controller and a Rule-Based (RB) Controller. The RB Controller showed a stronger impact than the NL Controller, which uses a small (7B/8B) LLM to simulate human-like interventions. Our findings indicate that agents particularly benefit from early interventions, leading to more efficient training and higher performance. Both intervention types outperform the baseline without interventions, highlighting the potential of LLM-mediated guidance to accelerate training and enhance MARL performance in challenging environments.

多智能体大模型强化学习干预引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。