arXiv:2608.16891cs.AIcs.CE2026-08

用可信执行层拦截危险操作,让AI的行动有安全边界。

Runtime Governance for Agentic AI: Action-Boundary Control with Trusted Provenance and Fail-Closed Execution

  • AI提出操作建议,由可信运行时决定是否执行
  • 测试中零风险操作被执行,所有操作都有可追溯记录
  • 适合需要严格控制自动化行为的高风险应用

Agentic AI系统会请求修改文件、发送消息、启动任务或改变工作流状态等操作,使安全问题从有害文本生成转向有害操作后果。提示层治理虽能影响模型行为,但无法建立执行边界。我们提出Aegis,一种运行时治理系统,将模型输出视为操作提案,通过可信决策层在工具执行前进行中介。模型提出;可信运行时决定。Aegis根据实时策略状态评估提案,服务端验证来源,不确定时默认拒绝,关键操作经类似参议院的共识机制审批,需多数同意并留有签名证据。我们在包含五个运行家族、42个任务、三种条件、每家族重复十次的沙箱数据集上评估Aegis。共6,300行提示-策略条件样本中出现79条高风险路径泄露;2,100行经Aegis治理的样本中,零次发生受控模拟工具调用,零次完成高风险副作用。全部1,832次尝试治理的操作均保留可信的Aegis溯源记录,全部1,019次参议院调解操作均有共识和最终签名计数证据。这些结果不证明通用自主代理安全,但支持更具体的结论:在此评估沙箱数据集中,运行时动作边界治理有效防止了已观察到的高风险提案演变为实际副作用。

原文摘要 · Abstract (English)

Agentic AI systems request tool actions that can modify files, send messages, launch jobs, or change workflow state. This shifts the safety problem from harmful text generation to harmful operational side effects. Prompt-level governance can shape model behavior, but it does not create an execution boundary. We introduce Aegis, a runtime governance system that treats model outputs as action proposals and mediates them through a trusted decision layer before tool execution. The model proposes; the trusted runtime decides. Aegis evaluates proposals against active policy state, resolves provenance server-side, fails closed under uncertainty, and routes selected cases through Senate-style settlement, a quorum- based non-unilateral authorization path. We evaluate Aegis on a repeated sandbox corpus spanning five run families, 42 tasks, three conditions, and ten repeats per family. Across 6,300 rows, prompt-policy conditioning produced 79 risky comparator-path leakage rows. Across 2,100 Aegis-governed rows, the system recorded zero governed mock-tool applications and zero governed risky side-effect completions. All 1,832 Aegis-attempted governed rows preserved trusted Aegis-resolved provenance, and all 1,019 Senate-settled rows had quorum and final signed tally evidence. These results do not prove general autonomous-agent safety. They support the narrower systems claim that, in this evaluated sandbox corpus, runtime action-boundary governance prevented observed risky proposals from becoming governed side effects.

AI安全运行时治理操作控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。