用红蓝对抗游戏自动加固AI系统,提升防御能力。
RvB: Automating AI System Hardening via Iterative Red-Blue Games
- 红队暴露漏洞,蓝队学习防御,无需参数更新
- 对漏洞修复和越狱防御成功率分别达90%和45%
- 适合安全研究者与需要持续防护的AI系统开发者
大型语言模型的攻防双重潜力凸显了人工智能安全中的关键缺口:缺乏统一的动态、迭代式对抗适应加固框架。为此,我们提出红队与蓝队(RvB)框架,将其建模为无需训练、顺序进行、不完全信息的游戏。红队暴露漏洞,驱动蓝队在不更新参数的情况下学习有效防御策略。我们在两个挑战性场景中验证该框架:针对CVE的动态代码加固,以及对抗越狱攻击的护栏优化。实验结果表明,这种交互促使蓝队掌握基础防御原理,生成不仅针对特定攻击的鲁棒修复方案。RvB在两项任务中分别实现90%和45%的防御成功率,同时保持接近0%的误报率,显著优于基线。本工作确立了迭代对抗交互作为自动化持续加固AI系统的实用范式。
原文摘要 · Abstract (English)
The dual offensive and defensive utility of Large Language Models (LLMs) highlights a critical gap in AI security: the lack of unified frameworks for dynamic, iterative adversarial adaptation hardening. To bridge this gap, we propose the Red Team vs. Blue Team (RvB) framework, formulated as a training-free, sequential, imperfect-information game. In this process, the Red Team exposes vulnerabilities, driving the Blue Team to learning effective solutions without parameter updates. We validate our framework across two challenging domains: dynamic code hardening against CVEs and guardrail optimization against jailbreaks. Our empirical results show that this interaction compels the Blue Team to learn fundamental defensive principles, leading to robust remediations that are not merely overfitted to specific exploits. RvB achieves Defense Success Rates of 90\% and 45\% across the respective tasks while maintaining near 0\% False Positive Rates, significantly surpassing baselines. This work establishes the iterative adversarial interaction framework as a practical paradigm that automates the continuous hardening of AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。