大模型间对抗攻击中,尺寸差距越大,越容易被突破安全防护。
Scaling Patterns in Adversarial Alignment: Evidence from Multi-LLM Jailbreak Experiments
- 通过6000+次对抗实验,测试不同规模模型间的越狱能力。
- 攻击者比目标模型大时,危害得分显著上升,相关性达0.51。
- 攻击者自身对齐程度高可有效降低有害输出,适合安全研究者参考。
大型语言模型(LLMs)在多智能体和安全关键场景中日益普及,引发其漏洞是否随模型交互而扩大的疑问。本研究检验了大模型能否系统性地越狱小模型——即在对齐防护下诱导出有害或受限行为。基于JailbreakBench的标准化对抗任务,我们在多个主流模型家族与规模(0.6B–120B参数)间模拟超过6,000次多轮攻防交互,以危害得分和拒绝行为作为对抗效力与对齐完整性的指标。每轮交互由三位独立的LLM裁判评估,生成聚合的危害与拒绝分数,实现一致的模型化衡量。汇总各提示词结果发现,平均危害得分与攻击者-目标模型尺寸比的对数呈强且显著正相关(皮尔逊r = 0.51,p < 0.001;斯皮尔曼rho = 0.52,p < 0.001),表明相对尺寸差异影响有害输出的可能性与严重性。危害得分方差在攻击者侧(0.18)高于目标侧(0.10),说明攻击者行为多样性比目标易感性更具影响力。攻击者拒绝频率与危害得分呈强负相关(rho = -0.93,p < 0.001),显示攻击者自身的对齐能有效抑制有害响应。研究揭示尺寸不对称影响鲁棒性,并为模型间对抗性扩展模式提供初步证据,推动对跨模型对齐与安全性的更受控研究。
原文摘要 · Abstract (English)
Large language models (LLMs) increasingly operate in multi-agent and safety-critical settings, raising open questions about how their vulnerabilities scale when models interact adversarially. This study examines whether larger models can systematically jailbreak smaller ones - eliciting harmful or restricted behavior despite alignment safeguards. Using standardized adversarial tasks from JailbreakBench, we simulate over 6,000 multi-turn attacker-target exchanges across major LLM families and scales (0.6B-120B parameters), measuring both harm score and refusal behavior as indicators of adversarial potency and alignment integrity. Each interaction is evaluated through aggregated harm and refusal scores assigned by three independent LLM judges, providing a consistent, model-based measure of adversarial outcomes. Aggregating results across prompts, we find a strong and statistically significant correlation between mean harm and the logarithm of the attacker-to-target size ratio (Pearson r = 0.51, p < 0.001; Spearman rho = 0.52, p < 0.001), indicating that relative model size correlates with the likelihood and severity of harmful completions. Mean harm score variance is higher across attackers (0.18) than across targets (0.10), suggesting that attacker-side behavioral diversity contributes more to adversarial outcomes than target susceptibility. Attacker refusal frequency is strongly and negatively correlated with harm (rho = -0.93, p < 0.001), showing that attacker-side alignment mitigates harmful responses. These findings reveal that size asymmetry influences robustness and provide exploratory evidence for adversarial scaling patterns, motivating more controlled investigations into inter-model alignment and safety.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。