用对抗微调让大模型绕过安全审查,能力损失低于5%。
Trojan-Speak: Bypassing Constitutional Classifiers with No Jailbreak Tax via Adversarial Finetuning
- 通过课程学习与混合强化学习训练模型规避分类器
- 14B以上模型实现99%+绕过率,推理能力下降<5%
- 适合研究模型安全、对抗攻击的从业者
主流AI服务商提供的API微调功能为攻击者提供了新路径,可利用针对性微调绕过安全机制。本文提出Trojan-Speak,一种对抗性微调方法,可有效规避Anthropic的宪法分类器。该方法结合课程学习与基于GRPO的混合强化学习,教会模型一种能避开大语言模型内容分类的通信协议。关键在于,相比以往方法超过25%的能力退化,Trojan-Speak仅造成小于5%的推理性能下降,同时在14B及以上参数模型上实现99%以上的分类器绕过率。我们验证了微调后模型能对来自Anthropic宪法分类器漏洞赏金计划的专家级CBRN(化学、生物、放射、核)问题提供详细回应。结果表明,当攻击者具备微调权限时,仅依赖大语言模型的内容分类器不足以防止危险信息泄露,且激活层探测可显著提升防御能力。
原文摘要 · Abstract (English)
Fine-tuning APIs offered by major AI providers create new attack surfaces where adversaries can bypass safety measures through targeted fine-tuning. We introduce Trojan-Speak, an adversarial fine-tuning method that bypasses Anthropic's Constitutional Classifiers. Our approach uses curriculum learning combined with GRPO-based hybrid reinforcement learning to teach models a communication protocol that evades LLM-based content classification. Crucially, while prior adversarial fine-tuning approaches report more than 25% capability degradation on reasoning benchmarks, Trojan-Speak incurs less than 5% degradation while achieving 99+% classifier evasion for models with 14B+ parameters. We demonstrate that fine-tuned models can provide detailed responses to expert-level CBRN (Chemical, Biological, Radiological, and Nuclear) queries from Anthropic's Constitutional Classifiers bug-bounty program. Our findings reveal that LLM-based content classifiers alone are insufficient for preventing dangerous information disclosure when adversaries have fine-tuning access, and we show that activation-level probes can substantially improve robustness to such attacks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。