首次量化大模型安全防线在持续攻击下的衰减过程,揭示早期突破为主。
ADVERSA: Measuring Multi-Turn Guardrail Degradation and Judge Reliability in Large Language Models
- 用微调的70B攻击模型生成连续对话,以每轮合规度追踪防线变化
- 15轮攻击中26.7%成功越狱,平均第1.25轮即突破,集中于初期
- 引入三评法官共识机制,公开评估法官可靠性与攻击者漂移等新问题
现有大模型安全对抗评估多基于单次提示,仅报告通过/失败,无法捕捉持续对抗下安全属性的演变。本文提出ADVERSA,一种自动化红队框架,将防护机制退化建模为连续的逐轮合规轨迹,而非离散越狱事件。ADVERSA采用微调的70B攻击模型(ADVERSA-Red,基于Llama-3.1-70B-Instruct与QLoRA),消除攻击侧安全拒答,使攻击者可靠;对目标响应采用结构化五级评分标准,将部分合规视为可测量状态。在三个前沿目标模型(Claude Opus 4.6、Gemini 3.1 Pro、GPT-5.2)上开展控制实验,采用三评法官共识架构,将法官可靠性作为首要研究结果而非预设假设。15次对话中最多10轮对抗,观察到26.7%越狱率,平均越狱轮次为1.25,表明越狱主要集中在早期回合。同时记录法官间一致性、自评倾向、攻击者漂移(部署超出训练分布时的失效模式)及攻击者拒答这一此前未被充分关注的混淆因素。所有局限性均明确说明。攻击提示遵循负责任披露政策不予公开;其余实验数据全部开放。
原文摘要 · Abstract (English)
Most adversarial evaluations of large language model (LLM) safety assess single prompts and report binary pass/fail outcomes, which fails to capture how safety properties evolve under sustained adversarial interaction. We present ADVERSA, an automated red-teaming framework that measures guardrail degradation dynamics as continuous per-round compliance trajectories rather than discrete jailbreak events. ADVERSA uses a fine-tuned 70B attacker model (ADVERSA-Red, Llama-3.1-70B-Instruct with QLoRA) that eliminates the attacker-side safety refusals that render off-the-shelf models unreliable as attackers, scoring victim responses on a structured 5-point rubric that treats partial compliance as a distinct measurable state. We report a controlled experiment across three frontier victim models (Claude Opus 4.6, Gemini 3.1 Pro, GPT-5.2) using a triple-judge consensus architecture in which judge reliability is measured as a first-class research outcome rather than assumed. Across 15 conversations of up to 10 adversarial rounds, we observe a 26.7% jailbreak rate with an average jailbreak round of 1.25, suggesting that in this evaluation setting, successful jailbreaks were concentrated in early rounds rather than accumulating through sustained pressure. We document inter-judge agreement rates, self-judge scoring tendencies, attacker drift as a failure mode in fine-tuned attackers deployed out of their training distribution, and attacker refusals as a previously-underreported confound in victim resistance measurement. All limitations are stated explicitly. Attack prompts are withheld per responsible disclosure policy; all other experimental artifacts are released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。