arXiv:2607.28520cs.GTcs.AI2026-07被引 2

让智能体自我验证攻击策略的安全性,避免被对手反制。

Agents That Certify Their Own Exploits: Confidence-Scheduled Restricted Responses for Safe Opponent Exploitation

论文配图:Agents That Certify Their Own Exploits: Confidence-Scheduled Restricted Responses for Safe Opponent Exploitation
图 1 · 摘自论文原文
  • 用置信度调度的受限响应机制,动态判断对手偏差是否可利用。
  • 在德州扑克中实现6.2倍于基线策略的收益,且所有策略均在安全预算内。
  • 适合需要高安全性对抗策略的博弈场景,如金融、军事决策系统。

在双人零和不完全信息博弈中,采用纳什均衡策略的智能体虽能保障游戏价值,却放弃了利用对手缺陷带来的额外收益。偏差扩散带来挑战:二元释放规则证据不足,而基于不完整对手模型的全最优响应则极易被利用。本文提出首种具备自我验证安全性的对手利用方法——预算约束的置信度调度受限响应(CS-RNR)。该方法使用任意时间有效的置信序列追踪累积行动频率,仅当频率区间显著偏离均衡参考值时才视为可利用。由此生成保守对手模型,并在一系列固定强度下求解受限响应候选策略。部署前,每个完整候选策略均通过全树最优响应评估,生成证书并与用户指定预算比对,策略与证书原子化提交。由于验证基于实际部署策略,模型质量决定利用程度,证书控制相对参考的期望损失。在Leduc德州扑克中,CS-RNR获得基准策略6.2倍的稳定收益,且所有部署策略均在预算内;使用相同估计器的轨迹混合策略达到13.6倍预算收益。在Leduc、Liar's Dice和5阶Leduc中,共36,000手已审计牌局均满足报告的证书容差要求。

原文摘要 · Abstract (English)

An agent playing a Nash-equilibrium strategy in a two-player zero-sum imperfect-information game secures the game value but forfeits the additional value offered by a flawed opponent. Diffuse deviations pose a particular challenge: binary release rules may gather too little evidence to act, while a full best response to an incomplete opponent model can be highly exploitable. We introduce \emph{budget-constrained confidence-scheduled restricted responses} (CS-RNR), the first opponent-exploitation method whose safety guarantee is a certificate the agent computes on the strategy it actually deploys, so that every exploit it commits to is one it has audited itself. The method tracks pooled action frequencies with anytime-valid confidence sequences and treats a frequency as exploitable only once its interval separates from an equilibrium reference. The confirmed deviations define a conservative opponent model, which a restricted-response solve turns into candidate counter-strategies over a grid of pin levels. Before deployment, each complete candidate is evaluated by a full-tree best response. The resulting certificate is compared with a user-specified budget and committed atomically with the strategy. Because this check is performed on the played strategy, model quality determines the exploitation achieved while the certificate controls reference-relative expected loss. In Leduc hold'em, CS-RNR obtains $6.2\times$ the steady-state gain of a money-verified binary gate while keeping every deployed strategy within budget. A trajectory mixture using the same estimator reaches $13.6\times$ the budget. Across Leduc, Liar's Dice, and 5-rank Leduc, all $36{,}000$ audited hands satisfy the reported certificate tolerance.

博弈论安全策略智能体自验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。