用大模型生成任务框架,让机器人团队更高效协作。
Generative World Models of Tasks: LLM-Driven Hierarchical Scaffolding for Embodied Agents
- 用大模型动态构建任务层级结构,替代盲目试错
- 在多智能体足球中提升样本效率,实现复杂策略学习
- 适合研究具身智能与高层决策的学者参考
当前智能体发展多依赖模型规模和交互数据的扩展,但在复杂长周期多智能体任务(如机器人足球)中,因探索空间过大且奖励稀疏,端到端方法常失效。我们提出,有效的世界模型需同时建模物理规律与任务语义。对2024年低资源多智能体足球研究的系统分析显示,符号化与层次化方法(如HTNs、BSNs)正与多智能体强化学习(MARL)融合,将复杂目标分解为可管理子目标,形成内在课程。为此,我们提出层级任务环境(HTE)框架,作为连接简单反应行为与高级团队策略的桥梁。该框架引入大语言模型(LLMs)作为任务的生成式世界模型,动态生成此结构化支撑。我们认为,HTE能引导探索、生成有意义学习信号,并使智能体内化层级结构,从而以更高样本效率训练出更通用、更强的智能体,优于纯端到端方法。
原文摘要 · Abstract (English)
Recent advances in agent development have focused on scaling model size and raw interaction data, mirroring successes in large language models. However, for complex, long-horizon multi-agent tasks such as robotic soccer, this end-to-end approach often fails due to intractable exploration spaces and sparse rewards. We propose that an effective world model for decision-making must model the world's physics and also its task semantics. A systematic review of 2024 research in low-resource multi-agent soccer reveals a clear trend towards integrating symbolic and hierarchical methods, such as Hierarchical Task Networks (HTNs) and Bayesian Strategy Networks (BSNs), with multi-agent reinforcement learning (MARL). These methods decompose complex goals into manageable subgoals, creating an intrinsic curriculum that shapes agent learning. We formalize this trend into a framework for Hierarchical Task Environments (HTEs), which are essential for bridging the gap between simple, reactive behaviors and sophisticated, strategic team play. Our framework incorporates the use of Large Language Models (LLMs) as generative world models of tasks, capable of dynamically generating this scaffolding. We argue that HTEs provide a mechanism to guide exploration, generate meaningful learning signals, and train agents to internalize hierarchical structure, enabling the development of more capable and general-purpose agents with greater sample efficiency than purely end-to-end approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。