用自然语言控制上千个智能体,实现人类与群体的高效协同。
Harnessing Language for Coordination: A Framework and Benchmark for LLM-Driven Multi-Agent Control
- 通过自然语言对话让单人指挥最多2000个智能体。
- 在实时策略游戏中完成移动协调、弱点利用等复杂任务。
- 适合研究人机协同、多智能体系统与LLM应用的学者。
大型语言模型在多种任务中表现出色,其在人类与大量智能体协作中的潜力尚未充分探索,但对灾难响应、城市规划和实时战略场景具有重要意义。本文提出(1)一个用于评估此类能力的实时策略游戏基准,以及(2)一种名为HIVE的新框架。HIVE使单个用户可通过与大语言模型的自然语言对话,协调多达2000个智能体。实验表明,该混合方法可有效完成智能体移动协调、利用单位弱点、结合人类标注信息,以及理解地形与战略点等任务。同时,研究也揭示当前模型在处理空间视觉信息和制定长期战略方面的显著局限。本工作为未来人-群协同研究提供了方向。HIVE项目页面(hive.syrkis.com)包含系统演示视频。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated remarkable performance across various tasks. Their potential to facilitate human coordination with many agents is a promising but largely under-explored area. Such capabilities would be helpful in disaster response, urban planning, and real-time strategy scenarios. In this work, we introduce (1) a real-time strategy game benchmark designed to evaluate these abilities and (2) a novel framework we term HIVE. HIVE empowers a single human to coordinate swarms of up to 2,000 agents through a natural language dialog with an LLM. We present promising results on this multi-agent benchmark, with our hybrid approach solving tasks such as coordinating agent movements, exploiting unit weaknesses, leveraging human annotations, and understanding terrain and strategic points. Our findings also highlight critical limitations of current models, including difficulties in processing spatial visual information and challenges in formulating long-term strategic plans. This work sheds light on the potential and limitations of LLMs in human-swarm coordination, paving the way for future research in this area. The HIVE project page, hive.syrkis.com, includes videos of the system in action.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。