用合成数据优化AI代理攻击,显著提升攻击强度。
Optimizing AI Agent Attacks With Synthetic Data
- 拆解攻击能力为五项技能,分别优化
- 合成数据训练下攻击安全分从0.87降至0.41
- 适合研究AI安全与对抗性测试的从业者
随着AI系统部署日益复杂且影响重大,评估其风险愈发重要。AI控制框架为此提供了一种方法,但有效评估需激发强攻击策略。在计算资源受限、真实数据稀缺的复杂代理环境中,这极具挑战。本文在SHADE-Arena这一多样化真实控制环境数据集上,将攻击能力分解为五项核心技能:怀疑建模、攻击选择、计划生成、执行与隐蔽性,并分别进行优化。为克服数据不足问题,我们构建了攻击动态的概率模型,通过仿真优化攻击超参数,再验证其在真实环境中的迁移效果。结果表明,该方法显著提升了攻击强度,使安全评分从基线0.87下降至0.41。
原文摘要 · Abstract (English)
As AI deployments become more complex and high-stakes, it becomes increasingly important to be able to estimate their risk. AI control is one framework for doing so. However, good control evaluations require eliciting strong attack policies. This can be challenging in complex agentic environments where compute constraints leave us data-poor. In this work, we show how to optimize attack policies in SHADE-Arena, a dataset of diverse realistic control environments. We do this by decomposing attack capability into five constituent skills -- suspicion modeling, attack selection, plan synthesis, execution and subtlety -- and optimizing each component individually. To get around the constraint of limited data, we develop a probabilistic model of attack dynamics, optimize our attack hyperparameters using this simulation, and then show that the results transfer to SHADE-Arena. This results in a substantial improvement in attack strength, reducing safety score from a baseline of 0.87 to 0.41 using our scaffold.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。