用博弈论设计AI安全机制,让监管更主动防骗。
Incentive-Aware AI Safety via Strategic Resource Allocation: A Stackelberg Security Games Perspective
- 将安全监管建模为攻防博弈,动态应对人为恶意
- 在资源有限时仍能有效防范数据污染和评估漏洞
- 适合关注制度设计与风险防御的研究者
随着AI系统日益自主,其安全可靠不仅依赖模型对齐,还需对开发与部署中的人类与机构进行战略监管。现有框架多将对齐视为静态优化问题,忽视了数据收集、模型评估与部署中动态的对抗性激励。本文提出基于斯塔克尔伯格安全博弈(Stackelberg Security Games, SSGs)的新视角:该博弈模型适用于不确定环境下的对抗性资源分配。将AI监管视为防御者(审计者、评估者、部署者)与攻击者(恶意行为者、偏离目标的贡献者或最坏情况故障模式)之间的战略互动,SSGs为全生命周期中的激励设计、监督能力有限性及对抗不确定性提供统一分析框架。我们展示该框架如何指导:(1) 训练阶段对抗数据/反馈污染的审计;(2) 评审资源受限下的预部署评估;(3) 对抗环境中多模型的稳健部署。该融合使算法对齐与制度设计相衔接,凸显博弈论威慑可使监管更具前瞻性、风险敏感性并抵御操纵。
原文摘要 · Abstract (English)
As AI systems grow more capable and autonomous, ensuring their safety and reliability requires not only model-level alignment but also strategic oversight of the humans and institutions involved in their development and deployment. Existing safety frameworks largely treat alignment as a static optimization problem (e.g., tuning models to desired behavior) while overlooking the dynamic, adversarial incentives that shape how data are collected, how models are evaluated, and how they are ultimately deployed. We propose a new perspective on AI safety grounded in Stackelberg Security Games (SSGs): a class of game-theoretic models designed for adversarial resource allocation under uncertainty. By viewing AI oversight as a strategic interaction between defenders (auditors, evaluators, and deployers) and attackers (malicious actors, misaligned contributors, or worst-case failure modes), SSGs provide a unifying framework for reasoning about incentive design, limited oversight capacity, and adversarial uncertainty across the AI lifecycle. We illustrate how this framework can inform (1) training-time auditing against data/feedback poisoning, (2) pre-deployment evaluation under constrained reviewer resources, and (3) robust multi-model deployment in adversarial environments. This synthesis bridges algorithmic alignment and institutional oversight design, highlighting how game-theoretic deterrence can make AI oversight proactive, risk-aware, and resilient to manipulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。