arXiv:2410.02511cs.AIcs.MA2024-10被引 12

用大模型识别关键状态,让多智能体探索更高效

Stop Wandering, Find the Keys: LLMs Discriminate Key States for Efficient Multi-Agent Exploration

  • 从大模型提取任务相关的关键状态,低成本精准引导
  • 设计奖励机制使智能体探索速度提升10倍,性能超越现有方法
  • 适合需要高效多智能体协同的复杂环境研究者

面对广阔的状-动空间,强化学习中的多智能体高效探索仍是长期难题。尽管新颖性、多样性与不确定性日益受关注,但缺乏有效引导的探索仍导致大量冗余消耗。本文提出LEMAE方法,利用知识型大语言模型(LLM)提供任务相关的有效指导,以符号化关键状态为核心,通过低开销的判别式方式将其落地。为激发关键状态潜力,设计基于子空间的回溯内在奖励(SHIR),提升其奖励密度以引导智能体聚焦;同时构建关键状态记忆树(KSMT),有序追踪任务中关键状态间的转移。得益于减少冗余探索,LEMAE在挑战性基准(如SMAC和MPE)上显著优于现有最先进方法,在特定场景下实现10倍加速。

原文摘要 · Abstract (English)

With expansive state-action spaces, efficient multi-agent exploration remains a longstanding challenge in reinforcement learning. Although pursuing novelty, diversity, or uncertainty attracts increasing attention, redundant efforts brought by exploration without proper guidance choices poses a practical issue for the community. This paper introduces a systematic approach, termed LEMAE, choosing to channel informative task-relevant guidance from a knowledgeable Large Language Model (LLM) for Efficient Multi-Agent Exploration. Specifically, we ground linguistic knowledge from LLM into symbolic key states, that are critical for task fulfillment, in a discriminative manner at low LLM inference costs. To unleash the power of key states, we design Subspace-based Hindsight Intrinsic Reward (SHIR) to guide agents toward key states by increasing reward density. Additionally, we build the Key State Memory Tree (KSMT) to track transitions between key states in a specific task for organized exploration. Benefiting from diminishing redundant explorations, LEMAE outperforms existing SOTA approaches on the challenging benchmarks (e.g., SMAC and MPE) by a large margin, achieving a 10x acceleration in certain scenarios.

多智能体强化学习大模型引导高效探索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。