用独立验证证明来管住自主AI的关键操作,不监控推理过程。
Governing Actions, Not Agents: Institutional Attestation as a Governance Model for Autonomous AI Systems
- 关键操作需独立权威方验证,与执行者分离
- 执行依赖多重可信证明,且记录可追溯
- 适合医疗、软件发布等高风险场景
自主AI系统可能执行具有重大后果且不可逆的操作,如临床开药和软件部署。本文观察到,人类机构治理强大自主行为体时,并非监控其推理过程,而是在关键行动节点要求独立证实的证据。我们形式化这一制度模式为一种面向AI系统的计算治理模型:代理保留完整的规划与推理自主权,但对高风险操作无执行权限;执行须满足由独立权威来源分别验证的预设条件,这些条件与声明意图密码绑定,并由确定性策略评估;决策记录于防篡改日志,支持独立复核。我们展示了概念验证实现,并以软件部署和临床开药为例说明该模型的应用。
原文摘要 · Abstract (English)
Autonomous AI agents may begin to perform consequential, irreversible actions such as clinical prescribing and production software deployment. This paper observes that human institutions have governed powerful autonomous actors not by monitoring their reasoning but by requiring independently attested evidence at the point of consequential action. We formalise this institutional pattern as a computational governance model for AI agent systems. Under the proposed model, an agent retains full autonomy over planning and reasoning but holds no execution authority over designated high-risk actions. Execution is conditional on preconditions that are each independently attested by a separate authoritative source, cryptographically bound to a declared intent, and evaluated by a deterministic policy. Decisions are recorded in a tamper-evident log amenable to independent re-verification. We present a proof-of-concept implementation and illustrate the model with examples from software deployment and clinical prescribing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。