设计防作弊保险合约,让自动驾驶AI诚实运行不钻空子。
Gaming-Resistant Insurance Contracts for Autonomous AI Agents: Strategy-Proof Toll Mechanism Design
- 用五类攻击场景分析保险合约漏洞,新增三类防御条款。
- 验证跨模型追踪数据中合约可抵抗策略性操作,保障安全执行。
- 适合研究自主智能体激励机制与可信保险设计的学者参考。
本文针对自主AI代理保险合约中的策略性行为,刻画了五种攻击方式,并证明了何种条件下可实现博弈抗性。已有方案通过最小权限和禁止拆分动作关闭两类攻击面,剩余三类需新条款:一是共控聚合防止跨边界重路由降低费用;二是将接口故障(如无效JSON)视为合同相关事件而非安全胜利,引入升级费逆转激励;三是采用基于组件最小惩罚的模型身份菜单,使真实报告部署模型成为弱占优策略。将这些条款与前作的运行时保证结合,实现全五类攻击空间下的联合激励相容。最终构建双参数保费族,在诚实均衡下满足个体理性与弱预算平衡。成果形成一套面向自主代理副作用的激励相容控制层。
原文摘要 · Abstract (English)
Paper A defines a time-consistent actuarial runtime that prices each side-effect-bearing action against a contractually fixed safe default and gates execution against a reserve budget. It treats the operator as passive. This paper makes the operator strategic. We characterise a five-attack space for autonomous AI-agent insurance contracts and prove when the actuarial runtime is gaming-resistant. Two attack surfaces -- post-toll safe-default selection and within-boundary action splitting -- are closed by Paper A's minimal-authority and no-splitting clauses. The remaining three require new contract clauses. First, common-control aggregation prevents cross-boundary re-routing from reducing toll below the boundary potential applied to total exposure. Second, interface failures such as invalid JSON are contract-relevant events, not safety wins: treating them as zero-toll safe defaults can reward unreliable models, while escalation fees reverse the incentive. We validate this interface-compliance theorem on committed cross-model traces from the companion empirical paper. Third, a model-identity menu with a componentwise-minimum penalty schedule makes truthful reporting of the deployed model weakly dominant. We then compose these clauses with Paper A's runtime guarantees to obtain joint incentive compatibility over the five-attack space. Finally, a two-parameter premium family discharges operator individual rationality and weak budget balance at the truthful equilibrium. The result is an incentive-compatibility layer for actuarial control of autonomous-agent side effects.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。