arXiv:2605.01415cs.AIcs.CY2026-05

AI安全核心是控制不可逆决策的扩散,而非仅保证输出正确。

AI Safety as Control of Irreversibility: A Systems Framework for Decision-Energy and Sovereignty Boundaries

  • 用决策能量密度量化系统中不可逆决策的集中风险。
  • 发现三大主权边界:决策权、资源调动权、自我扩展权,决定AI是否失控。
  • 提出边界稳定定理:只需防止高效节点释放不可逆权力,无需证明系统永远正确。

当前AI系统的能力增长与部署之间距离缩短。以往高风险技术受限于资本强度、物理瓶颈、组织惰性及专业供应链,而AI能力可低成本复制、嵌入流程并跨机构扩展。本文认为,部署摩擦降低从根本上改变了安全问题。安全不仅是局部输出正确或偏好对齐,更是对高决策密度下不可逆性的控制。论文引入决策能量密度概念,衡量节点生成、评估、选择和执行关键决策的能力速率。识别出三类主权边界:不可逆决策权、物理资源动员权、自我扩展权。模型显示,效率压力、路径依赖、规模反馈与弱边界约束会使决策能量集中在最高效的节点,导致责任分散,即使单次动作错误率低,仍可能引发系统级不可逆损失。主要成果为边界稳定定理:安全不需证明先进系统始终正确,而是需设计制度与技术机制,阻止单一高效率节点释放不可逆权力。论文将AI安全重构为分层控制、授权与外部可审查的限制,关联对齐、安全工程、组织经济学与制度设计。

原文摘要 · Abstract (English)

Recent AI systems compress the distance between capability growth and capability deployment. Earlier high-risk technologies were slowed by capital intensity, physical bottlenecks, organizational inertia, and specialized supply chains. By contrast, AI capabilities can be copied, invoked, embedded in workflows, and scaled across institutions at low marginal cost. This paper argues that declining deployment friction changes the safety problem at its root. Safety is not only local output correctness or preference alignment, but the control of irreversibility under rising decision density. The paper formalizes this claim through decision-energy density: the rate-weighted capacity of a node to generate, evaluate, select, and execute consequential decisions. It then identifies three sovereignty boundaries that determine whether AI remains an amplifier within a human-governed system or becomes a de facto control center: irreversible decision authority, physical resource mobilization authority, and self-expansion authority. The model shows how efficiency pressure, path dependence, scale feedback, and weak boundary constraints concentrate decision-energy in the most efficient node. This concentration can diffuse responsibility and raise the probability of irreversible system-level loss even when local per-action error rates remain low. The main result is a boundary stabilization theorem. It shows that safety need not require proving that advanced systems are always correct. Instead, it requires institutional and technical designs that prevent irreversible power from being released by a single high-efficiency node. The paper reframes AI safety as layered control, authorization, and externally reviewable limits, linking alignment, security engineering, organizational economics, and institutional design.

AI安全系统控制不可逆性边界设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。