构建可视化平台评估智能体在复杂事件中的协作能力
A Visualized Framework for Event Cooperation with Generative Agents
- 设计地图编辑器与动画模拟器,实现物理环境的可视化交互
- 提出8类事件测试集,包含基础与高难度变体,评估智能体表现
- 发现大模型在复杂协作中存在明显短板,适合研究多智能体系统者参考
大型语言模型(LLMs)已推动智能体社会的仿真发展,实现自主规划、记忆形成与社交互动。然而,现有框架常缺乏对事件组织的系统评估,且缺少与物理环境的可视化整合,限制了智能体在空间导航和物品交互上的真实感。本文开发了MiniAgentPro可视化平台,配备直观的地图编辑器用于自定义环境,并集成具备平滑动画的模拟播放器。基于该工具,构建了一个包含八种多样化事件场景的综合测试集,涵盖基础与高难度变体,用于评估智能体能力。使用GPT-4o进行评估显示,在基础设置中表现良好,但在高难度变体中暴露出明显的协作挑战。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have revolutionized the simulation of agent societies, enabling autonomous planning, memory formation, and social interactions. However, existing frameworks often overlook systematic evaluations for event organization and lack visualized integration with physically grounded environments, limiting agents' ability to navigate spaces and interact with items realistically. We develop MiniAgentPro, a visualization platform featuring an intuitive map editor for customizing environments and a simulation player with smooth animations. Based on this tool, we introduce a comprehensive test set comprising eight diverse event scenarios with basic and hard variants to assess agents' ability. Evaluations using GPT-4o demonstrate strong performance in basic settings but highlight coordination challenges in hard variants.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。