arXiv:2606.23991cs.AIcs.LG2026-06被引 1

厘清AI代理与自主性的本质差异,提出可自我演化的通用智能体架构。

Critique of Agent Model

论文配图:Critique of Agent Model
图 1 · 摘自论文原文
  • 从目标、身份、决策等五维度定义真正自主的智能体系统。
  • 提出GIC架构,支持自适应目标分解与基于世界模型的模拟推理。
  • 强调高自主系统仍需人类可控,为安全部署提供设计思路。

随着大型语言模型被宣传为“编程代理”“AI合作者”等“代理型”工具,以及关于机器自主性可能脱离人类控制的潜在威胁,明确自动化与真正自主性的边界变得至关重要。本文借鉴笛卡尔对独立思考的论述及科幻作品中对自主生命体的刻画,分析当前AI代理架构,提出真正的自主性需在系统内部实现目标、身份、决策、自我调节和学习等结构的内化,而非依赖外部流程搭建。由此区分‘代理型’(能力源于工程工作流)与‘自主型’(能力内生演化)系统,界定任务导向与开放世界自主运行的分界线。基于此,提出通用智能体架构GIC,融合层级目标分解、身份演化、基于独立训练世界模型的模拟推理、学习型自我调节及真实与模拟经验驱动的自主学习。同时探讨具备更高自主性的系统在可审计性、可控制性与安全性方面的关键挑战与应对策略。

原文摘要 · Abstract (English)

What is an agent? What constitutes agency? With the rise of Large Language Model (LLM) systems marketed as ``coding agents'', ``AI co-scientists'', and other ``agentic" tools that promise to drive up productivity, and at the same time, ``existential" concerns such as AI escaping human control with destructive power under a speculative ``machine agency" against humans, it has become essential to clarify where automation ends and agency begins, both for building capable systems and for understanding whether and what to fear. Drawing on Descartes' grounding of agency in independent thought, and on portrayals of autonomous beings in science fiction, we survey the current landscape of AI agents, and analyze agent architectures along five dimensions: goal, identity, decision-making, self-regulation, and learning. Specifically, we argue that genuine agency requires these structures to be \emph{internalized within the system itself} rather than assembled through external scaffolding. This distinction between \emph{agentic} systems, whose competence resides in engineered workflows, and \emph{agentive} systems, whose capabilities (including social interaction) arise endogenously, defines the boundary between systems designed for prescribed tasks, and those capable of operating in the open world with true autonomy. Building on this analysis, we propose the Goal-Identity-Configurator (GIC) architecture for a general-purpose agent model, combining hierarchical goal decomposition, identity evolution, simulative reasoning grounded in a separately trained world model, learned self-regulation, and self-directed learning from both real and simulated experience. Furthermore, we share insight on the auditability, controllability, and safety of agentive systems that possess greater autonomy and ``agency", but remain under human oversight.

智能体自主性架构设计安全可控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。