提出可信赖的随机策略验证框架,解决AI代理在不确定环境下的安全监控难题。
Efficient and Sound Probabilistic Verification for AI Agents

- 基于分布鲁棒优化计算违规概率上界,无需独立性假设
- 在终端与工具调用代理基准测试中,违规概率上界更紧致且效率更高
- 适合需要严格安全约束的高风险场景,如隐私保护系统
在复杂数字环境中运行的AI代理的安全保障已成为关键需求。基于形式语言(如Datalog)表达并执行策略的运行时监控方法具有前景,但现有方法仅适用于确定性策略。在实际应用中,面对不确定性,常需处理具有失败概率的随机谓词或状态转移(如存在误判可能的去标识化器或个人身份信息(PII)检测器)。此外,许多场景难以满足先前工作所需的独立性假设。本文提出一种基于分布鲁棒优化的可靠且高效的验证框架,可在不依赖预测间独立性的前提下,计算出政策违规概率的可靠上界。在标准的终端和工具调用代理基准测试中,所提方法优于现有技术,在保障严格违规概率边界的同时,提升了安全与实用性的平衡。
原文摘要 · Abstract (English)
Securing AI agents that operate in complex digital environments has become a critical need, and runtime monitoring approaches that formulate and enforce policies expressed in a formal language like Datalog offer a promising solution. However, existing approaches are restricted to deterministic policies. In many practical applications of AI agents, there is a need to enforce security policies in the face of ambiguity, leading to probabilistic predicates or state transitions (for example, a declassifier or Personally Identifiable Information (PII) detector that has some failure probability on each invocation). Furthermore, in many such applications, one cannot easily make the independence assumptions necessary to invoke prior work on probabilistic inference in Datalog. We address this by introducing a sound and efficient framework for such verification based on distributionally robust optimization, computing sound upper bounds on the probability of policy violation regardless of possible correlations between predicates. On standard benchmarks for terminal and tool calling agents, we demonstrate that our approach outperforms prior art and improves the security-utility trade-off while ensuring rigorous bounds on the probability of policy violation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。