为自主人工智能设计可量化的等级体系,判断其是否接近通用智能。
An Operational Kardashev-Style Scale for Autonomous AI - Towards AGI and Superintelligence
- 构建十维能力维度,通过加权几何平均生成综合指数
- 提出可测量的自我提升系数κ,实现自进化AI的可验证评估
- 适用于评估当前AI系统进展,适合关注AGI发展的研究者
我们提出一个受卡达舍夫文明等级启发但具备操作性的自主人工智能(AAI)等级体系,用于衡量从固定机器人流程自动化(AAI-0)到完全的人工通用智能(AAI-4)及更高级别的发展进程。该体系为多轴且可测试,定义了十个能力维度(自主性、泛化性、规划、记忆/持续性、工具经济性、自我修正、社会性/协调性、具身性、世界模型保真度、经济吞吐量),并通过复合的AAI指数(加权几何平均)聚合。引入可测量的自我改进系数κ(单位代理资源引发的能力增长),并定义维护与扩展两种封闭性属性,使“自进化AI”成为可证伪的标准。提出OWA-Bench——一个开放世界代理评估套件,用于评测长周期、工具使用、持续性智能体。通过各轴阈值、κ值和封闭性证明定义了AAI-0至AAI-4的层级门槛。合成实验展示当前系统在该尺度上的定位,以及随着自进化能力提升,委托性前沿(质量与自主性平衡)如何演进。还证明一个定理:在充分条件下,AAI-3智能体将随时间演变为AAI-5,形式化了‘婴儿级AGI’走向超智能的直觉。
原文摘要 · Abstract (English)
We propose a Kardashev-inspired yet operational Autonomous AI (AAI) Scale that measures the progression from fixed robotic process automation (AAI-0) to full artificial general intelligence (AAI-4) and beyond. Unlike narrative ladders, our scale is multi-axis and testable. We define ten capability axes (Autonomy, Generality, Planning, Memory/Persistence, Tool Economy, Self-Revision, Sociality/Coordination, Embodiment, World-Model Fidelity, Economic Throughput) aggregated by a composite AAI-Index (a weighted geometric mean). We introduce a measurable Self-Improvement Coefficient $κ$ (capability growth per unit of agent-initiated resources) and two closure properties (maintenance and expansion) that convert ``self-improving AI'' into falsifiable criteria. We specify OWA-Bench, an open-world agency benchmark suite that evaluates long-horizon, tool-using, persistent agents. We define level gates for AAI-0\ldots AAI-4 using thresholds on the axes, $κ$, and closure proofs. Synthetic experiments illustrate how present-day systems map onto the scale and how the delegability frontier (quality vs.\ autonomy) advances with self-improvement. We also prove a theorem that AAI-3 agent becomes AAI-5 over time with sufficient conditions, formalizing "baby AGI" becomes Superintelligence intuition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。