研究大模型红队攻击中能力差距的影响,发现攻击成功率随目标模型更强而急剧下降。
Capability-Based Scaling Trends for LLM-Based Red-Teaming
- 用能力差衡量攻防双方实力,模拟人类红队进行多组攻击测试
- 当目标模型能力超过攻击者时,攻击成功率骤降,且与MMLU-Pro社会学科表现强相关
- 提出可预测攻击成功率的缩放曲线,警示固定能力攻击者将失效
随着大语言模型能力与自主性的提升,通过红队测试识别漏洞对安全部署愈发重要。传统提示工程方法在‘弱攻强’场景下可能失效,即目标模型能力超越红队攻击者。为此,我们从攻击者与目标间的能力差距视角研究红队问题。通过600多组基于LLM的越狱攻击实验,覆盖多种模型家族、规模与能力层级,发现三大趋势:(i) 更强模型作为攻击者表现更优;(ii) 当目标模型能力超过攻击者时,攻击成功率显著下降;(iii) 攻击成功率与MMLU-Pro社会科学子集的高分表现高度相关。据此推导出基于能力差距的越狱缩放曲线,可预测特定目标下的攻击成功率。结果表明,固定能力攻击者(如人类)未来将失效,更具能力的开源模型会加剧现有系统风险,模型提供方需精准测量并控制模型的说服与操控能力以限制其攻击效力。
原文摘要 · Abstract (English)
As large language models grow in capability and agency, identifying vulnerabilities through red-teaming becomes vital for safe deployment. However, traditional prompt-engineering approaches may prove ineffective once red-teaming turns into a \emph{weak-to-strong} problem, where target models surpass red-teamers in capabilities. To study this shift, we frame red-teaming through the lens of the \emph{capability gap} between attacker and target. We evaluate more than 600 attacker-target pairs using LLM-based jailbreak attacks that mimic human red-teamers across diverse families, sizes, and capability levels. Three strong trends emerge: (i) more capable models are better attackers, (ii) attack success drops sharply once the target's capability exceeds the attacker's, and (iii) attack success rates correlate with high performance on social science splits of the MMLU-Pro benchmark. From these observations, we derive a \emph{jailbreaking scaling curve} that predicts attack success for a fixed target based on attacker-target capability gap. These findings suggest that fixed-capability attackers (e.g., humans) may become ineffective against future models, increasingly capable open-source models amplify risks for existing systems, and model providers must accurately measure and control models' persuasive and manipulative abilities to limit their effectiveness as attackers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。