提出45度定律与因果信任阶梯,系统规划可信AGI发展路径
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI
- 构建因果信任阶梯框架,分三层应对AGI安全挑战
- 定义感知到协作共五级可信度,逐层推进可信AGI建设
- 为高自主性系统提供可落地的安全治理路线图
确保人工智能通用智能(AGI)可靠避免有害行为是关键挑战,尤其在高自主性或安全敏感领域。尽管已有多种安全保障方案和极端风险警示,但平衡AI安全与能力的全面指南仍缺失。本文提出'AI-45°律'作为可信AGI发展的指导原则,并引入'可信AGI因果阶梯'作为实践框架。该框架借鉴朱迪亚·珀尔的因果之梯思想,系统化地分类当前AI能力与安全研究,包含近似对齐层、可干预层和可反思层三个核心层级,分别应对AGI及现代AI系统中的安全与可信性关键问题。基于此框架,我们定义了五个可信度层级:感知、推理、决策、自主性和协作可信度,代表可信AGI的不同且递进的方面。最后,我们提出一系列潜在治理措施以支持可信AGI的发展。
原文摘要 · Abstract (English)
Ensuring Artificial General Intelligence (AGI) reliably avoids harmful behaviors is a critical challenge, especially for systems with high autonomy or in safety-critical domains. Despite various safety assurance proposals and extreme risk warnings, comprehensive guidelines balancing AI safety and capability remain lacking. In this position paper, we propose the \textit{AI-\textbf{$45^{\circ}$} Law} as a guiding principle for a balanced roadmap toward trustworthy AGI, and introduce the \textit{Causal Ladder of Trustworthy AGI} as a practical framework. This framework provides a systematic taxonomy and hierarchical structure for current AI capability and safety research, inspired by Judea Pearl's ``Ladder of Causation''. The Causal Ladder comprises three core layers: the Approximate Alignment Layer, the Intervenable Layer, and the Reflectable Layer. These layers address the key challenges of safety and trustworthiness in AGI and contemporary AI systems. Building upon this framework, we define five levels of trustworthy AGI: perception, reasoning, decision-making, autonomy, and collaboration trustworthiness. These levels represent distinct yet progressive aspects of trustworthy AGI. Finally, we present a series of potential governance measures to support the development of trustworthy AGI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。