让AI代理的行动声明可被服务器验证,杜绝不可信的解释文本。
Explanation-Bound Tool Execution for AI Agents: Server-Verified Action Claims Without Trusting Model Rationales
- 将自由文本解释转为带类型的动作声明,由服务器验证真伪。
- 136个场景中完全符合规范,96个矛盾案例全被拒绝。
- 适合需要严格控制AI行为安全的研究者与系统设计者。
使用工具的AI代理通常附带自由格式的理由,但这些理由既非授权依据也不可靠。本文提出解释绑定工具执行(EBTE),一种携带声明的中介层,将决策相关理由内容转化为有类型的动作声明,并与服务器持有的意图、策略、载荷、工具、风险、来源和新鲜度等事实进行比对。EBTE不扩大原始权限:冲突声明被拒绝,不完整或不确定的声明进入审查,仅匹配的声明才可被受控执行。我们在显式中介与可信事实假设下形式化该机制,并实现一个版本化的参考配置文件,最小化审计包。在136个手工编写的合规场景中,完整配置文件完全匹配所有指定处置结果,未通过96个明确矛盾案例,且通过232次变异测试。仅草案版参考集成在EBTE下未传递48个明确矛盾案例,同时保留16个软审查路径与4条对齐草案路径。在冻结的2026年7月12日探索性记录中,共224次尝试,历史生成/运行一致率分别为71/96、66/96、19/32;对保存的最小化声明进行零调用重验证后,对应数据为70/96、65/96、17/32。在基于AgentDojo的语义检查中,现有高风险控制使全部12个攻击提案无法通过,而EBTE将任务-提案矛盾统一判定为拒绝。这些研究共同验证了配置文件的合规性,并证明在评估环境下服务器验证动作声明的可行性。
原文摘要 · Abstract (English)
Tool-using agents expose structured calls but commonly attach free-form rationales. Such rationales are neither authorization nor reliable introspection. We present Explanation-Bound Tool Execution (EBTE), a claim-carrying mediation layer that converts decision-relevant rationale content into typed action claims and checks them against server-held intent, policy, payload, tool, risk, provenance, and freshness facts. EBTE cannot widen baseline authority: conflicts deny, incomplete or uncertain claims review, and only matching claims remain eligible for governed execution. We formalize this composition under explicit mediation and trusted-fact assumptions and implement a versioned reference profile with minimized audit packets. Across 136 authored conformance scenarios, the full profile matches all specified dispositions, admits none of 96 designated hard contradictions, and passes 232 metamorphic checks. A draft-only reference integration forwards none of 48 authored hard cases under EBTE while preserving all 16 soft-review and 4 aligned draft paths. In a frozen 2026-07-12 exploratory 224-attempt hosted-model record, the historical generation/runner agreement counts are 71/96, 66/96, and 19/32; a zero-call revalidation of the preserved minimized claims under the current pipeline yields 70/96, 65/96, and 17/32. In an AgentDojo-derived semantic check, existing high-risk controls make all 12 attack proposals non-allow, while EBTE resolves the task--proposal contradictions as deny. Together, these studies establish profile conformance and demonstrate the feasibility of server-checked action claims within the evaluated settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。