提出统一心智模型,让大模型智能体具备人类级认知能力。
Unified Mind Model: Reimagining Autonomous Agents in the LLM Era
- 基于全局工作空间理论,融合大模型构建多模态感知与推理架构。
- 实现规划、记忆、反思、工具使用等人类级认知功能。
- 开发无需编程的MindOS引擎,快速生成特定任务智能体。
大语言模型(LLMs)在多个领域、任务和语言上展现出卓越能力(如ChatGPT和GPT-4),重新激发了对具备类人认知能力的通用自主智能体的研究。这类人类级智能体需要语义理解与指令遵循能力,而这正是LLMs的核心优势。尽管已有若干基于LLM的人类级智能体初步尝试,但其理论基础仍是一个重大开放问题。本文提出一种新型理论认知架构——统一心智模型(UMM),为构建具备人类级认知能力的自主智能体提供指导。具体而言,我们的UMM以全局工作空间理论为基础,进一步利用LLMs赋予智能体多模态感知、规划、推理、工具使用、学习、记忆、反思与动机等多种认知能力。基于UMM,我们进一步开发了智能体构建引擎MindOS,使用户无需任何编程即可快速创建领域或任务特定的自主智能体。
原文摘要 · Abstract (English)
Large language models (LLMs) have recently demonstrated remarkable capabilities across domains, tasks, and languages (e.g., ChatGPT and GPT-4), reviving the research of general autonomous agents with human-like cognitive abilities. Such human-level agents require semantic comprehension and instruction-following capabilities, which exactly fall into the strengths of LLMs. Although there have been several initial attempts to build human-level agents based on LLMs, the theoretical foundation remains a challenging open problem. In this paper, we propose a novel theoretical cognitive architecture, the Unified Mind Model (UMM), which offers guidance to facilitate the rapid creation of autonomous agents with human-level cognitive abilities. Specifically, our UMM starts with the global workspace theory and further leverage LLMs to enable the agent with various cognitive abilities, such as multi-modal perception, planning, reasoning, tool use, learning, memory, reflection and motivation. Building upon UMM, we then develop an agent-building engine, MindOS, which allows users to quickly create domain-/task-specific autonomous agents without any programming effort.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。