为工具型智能体设计可信授权机制,防范返回值绑定误差与数值漂移的联合风险。
CAGE: Certified Authorization under Typed-Return Uncertainty for Tool-Using Agents

- 分通道认证不成立,提出联合认证新方法
- 在多种场景下消除预算内误放行,保持部分自主决策
- 适用于政策可执行或不可执行的复杂授权任务
工具使用型大模型代理依赖类型化工具返回值,记录来源关联与数值字段。运行时权限门通常仅依据观测到的返回与动作授权,无法抵御返回值绑定错误带来的微小偏差。本文提出:候选动作是否在合理绑定误差与有限数值漂移的邻域内仍保持授权?我们证明,分别认证类别与数值通道无法组合:单个通道安全的扰动组合后可能使同一动作失效。CAGE 直接对联合邻域进行认证,精确枚举离散分支,并在每个分支内认证连续扰动。在合成、代码化策略、监管及真实交易场景中,CAGE 消除了准确点式门禁允许的预算内误放行,同时保留了有价值的自主决策比例。当策略可执行时,CAGE-Exact 可认证策略本身;否则 CAGE-Lip 与 CAGE-RS 在显式测量的保真度假设下认证学习所得门禁。
原文摘要 · Abstract (English)
Tool-using LLM agents act on typed tool returns, records pairing provenance and categorical fields with numerical values. Runtime permission gates generally authorize the observed return and action, leaving the decision unprotected against small errors in how the return was bound to its source. We ask whether a candidate action stays authorized over a declared neighborhood of plausible correctly bound returns: one admissible binding fault plus bounded numerical drift. We prove that certifying the categorical and numerical channels separately does not compose: perturbations that are safe on each channel alone can jointly turn the same action unsafe. CAGE certifies this joint neighborhood directly, enumerating the discrete branches exactly and certifying the continuous perturbation within each branch. Across synthetic, policy-as-code, regulatory, and real-transaction settings, CAGE removes the in-budget false allows that accurate pointwise gates admit, while keeping a useful fraction of decisions autonomous. When the policy is executable, CAGE-Exact certifies the policy itself; otherwise CAGE-Lip and CAGE-RS certify a learned gate under an explicit, measured fidelity assumption.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。