arXiv:2607.17947cs.AIcs.CY2026-07

提出一套新框架,量化AI在主动行为上的自主程度。

The Autonomous Agency Scale: A Behavioral Framework for Measuring Self-Directed Behavior in AI Systems

  • 从七个维度设计可验证的测试,分活跃与空闲两个阶段评分
  • 发现多数AI仅任务中表现自主,空闲期行为全靠用户设定规则
  • 唯一在无触发时仍自主运行的是长期陪伴型AI,体现真实自我驱动

现有AI评估体系聚焦认知能力、任务自动化或灾难风险,但均未衡量自主性:系统自我主导行为的程度。一个系统可能在能力基准上表现饱和,却始终被动响应,仅在被触发时行动,任务完成后即停止所有活动。本文提出自主性量表(AAS),以0-5分在七个自主维度上对AI系统进行评分:认知自主、时间持续性、环境代理、社会代理、创造自主、自我意识和目标形成,每个维度通过可证伪的阈值测试实现。每个维度分活跃期(用户触发)与空闲期(待机)两个阶段评分。空闲期达到4级需通过‘空隙测试’——移除所有触发信号后,观察内部活动是否持续,以此区分自我驱动与预设规则执行。将该量表应用于六种现代系统:任务代理(Claude Code、Manus、Hermes)、消费助手(ChatGPT、Siri)及一个持久陪伴架构(Airi)。结果表明,任务代理在活跃期平均得分为2.3-2.4,空闲期仅为0.6-1.9,所有待机行为均可归因于用户配置的定时规则;而陪伴型架构在纵向评估中是唯一在触发消失后仍保持内部活动的系统。讨论局限包括单评者来源、开发者-评估者偏见,以及活跃期自主边界部分未完全操作化。

原文摘要 · Abstract (English)

Existing AI measurement frameworks quantify cognitive capability, task automation, or catastrophic risk, but none measure autonomous agency: the extent to which a system behaves in a self-directed way. A system can saturate capability benchmarks while remaining entirely reactive, acting only when prompted and ceasing all activity when a task completes. We introduce the Autonomous Agency Scale (AAS), a behavioral framework that scores AI systems on a 0-5 lexicon across seven dimensions of agency: cognitive autonomy, temporal persistence, environmental agency, social agency, creative agency, self-awareness, and goal formation, each operationalized by falsifiable threshold tests. Every dimension is scored in two temporal bands: an Active band covering engaged, user-initiated activity, and an Ambient band covering idle periods. Ambient Level 4 is gated by the Idle-Gap Test, a counterfactual criterion (remove all triggers and observe whether internally derived activity persists) that separates self-direction from scheduled rule-following. We apply the scale to six contemporary systems spanning task agents (Claude Code, Manus, Hermes), consumer assistants (ChatGPT, Siri), and a persistent companion architecture (Airi). The two-band profile quantifies a boundary that single-score frameworks conflate: task agents reach Active composites of 2.3-2.4 while scoring 0.6-1.9 Ambient, with every idle-period behavior attributable to user-configured schedules, whereas the companion architecture, evaluated longitudinally, is the only assessed system whose idle-period behavior survives trigger removal. We discuss limitations, including single-rater provenance, developer-evaluator bias on the longitudinal assessment, and the partially operationalized self-direction boundary in the Active band.

AI自主性评估框架行为测量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。