厘清前沿AI中对齐、代理与自主性的定义分歧,为安全治理提供系统视角。
Alignment, Agency and Autonomy in Frontier AI: A Systems Engineering Perspective
- 从多学科角度梳理对齐、代理与自主性的演变,揭示定义差异的根源。
- 通过特斯拉、波音等案例,揭示自主系统误判带来的真实风险。
- 适合关注AI治理、安全设计与高阶应用的开发者与政策制定者。
随着人工智能的规模扩展,对齐、代理性和自主性已成为AI安全、治理与控制的核心议题。然而,这些概念在人类语境中尚无统一定义,不同学科(如哲学、心理学、法律、计算机科学等)存在显著差异,导致其在人工智能领域应用时产生理解分歧,进而引发系统设计与监管策略的冲突。本文追溯这些概念的历史、哲学与技术演进,强调其定义如何影响AI的研发、部署与监督。我们指出,对对齐与自主性的紧迫关注不仅源于技术进步,更因AI日益参与高风险决策。以代理型AI为例,分析机器代理性与自主性的涌现特性,凸显现实系统中对齐失效的风险。通过剖析自动化失败案例(特斯拉Autopilot、波音737 MAX)、多智能体协作(Meta CICERO)及演化架构(DeepMind AlphaZero、OpenAI AutoGPT),评估前沿AI带来的治理与安全挑战。
原文摘要 · Abstract (English)
As artificial intelligence scales, the concepts of alignment, agency, and autonomy have become central to AI safety, governance, and control. However, even in human contexts, these terms lack universal definitions, varying across disciplines such as philosophy, psychology, law, computer science, mathematics, and political science. This inconsistency complicates their application to AI, where differing interpretations lead to conflicting approaches in system design and regulation. This paper traces the historical, philosophical, and technical evolution of these concepts, emphasizing how their definitions influence AI development, deployment, and oversight. We argue that the urgency surrounding AI alignment and autonomy stems not only from technical advancements but also from the increasing deployment of AI in high-stakes decision making. Using Agentic AI as a case study, we examine the emergent properties of machine agency and autonomy, highlighting the risks of misalignment in real-world systems. Through an analysis of automation failures (Tesla Autopilot, Boeing 737 MAX), multi-agent coordination (Metas CICERO), and evolving AI architectures (DeepMinds AlphaZero, OpenAIs AutoGPT), we assess the governance and safety challenges posed by frontier AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。