用控制理论重构AI对齐,构建分层可互操作的对齐框架
Out of Control -- Why Alignment Needs Formal Control Theory (and an Alignment Control Stack)
- 将对齐问题重述为形式化最优控制,建立分层控制架构
- 提出对齐控制栈,明确各层测量与控制特性及形式化接口
- 适合关注安全监管与系统可靠性的研究者和政策制定者
本文主张将形式化最优控制理论作为人工智能对齐研究的核心,提供不同于当前安全与安全方法的新视角。尽管近期在安全性和机制可解释性方面的研究推进了对齐的形式化方法,但这些方法往往难以满足其他技术领域控制框架所需的泛化能力。此外,缺乏关于如何使不同对齐/控制协议实现互操作的研究。我们提出,通过以形式化最优控制原则重构对齐,并根据物理到社会技术层面的层级结构定义控制应用方式,可以更好地理解前沿模型与代理型AI系统的控制潜力与局限。为此,我们引入对齐控制栈,明确各层级的测量与控制特征,以及各层间的正式互操作性。我们认为,此类分析对于政府和监管机构建立可持续的保障体系至关重要。最终目标是将经过验证的控制理论与实际部署需求相结合,构建更全面的对齐框架,提升先进AI系统在安全与可靠性方面的方法。
原文摘要 · Abstract (English)
This position paper argues that formal optimal control theory should be central to AI alignment research, offering a distinct perspective from prevailing AI safety and security approaches. While recent work in AI safety and mechanistic interpretability has advanced formal methods for alignment, they often fall short of the generalisation required of control frameworks for other technologies. There is also a lack of research into how to render different alignment/control protocols interoperable. We argue that by recasting alignment through principles of formal optimal control and framing alignment in terms of hierarchical stack from physical to socio-technical layers according to which controls may be applied we can develop a better understanding of the potential and limitations for controlling frontier models and agentic AI systems. To this end, we introduce an Alignment Control Stack which sets out a hierarchical layered alignment stack, identifying measurement and control characteristics at each layer and how different layers are formally interoperable. We argue that such analysis is also key to the assurances that will be needed by governments and regulators in order to see AI technologies sustainably benefit the community. Our position is that doing so will bridge the well-established and empirically validated methods of optimal control with practical deployment considerations to create a more comprehensive alignment framework, enhancing how we approach safety and reliability for advanced AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。