用对抗推理提升越狱成功率,发现大模型新漏洞
Adversarial Reasoning at Jailbreaking Time
- 通过损失信号引导推理时计算,实现自动越狱
- 在多个对齐模型上达到当前最高攻击成功率
- 为构建更鲁棒的AI系统提供新思路
随着大语言模型能力增强和广泛应用,研究其失效案例变得日益重要。近期在测试时计算资源标准化、度量与扩展方面的进展,为优化模型在高难度任务上的表现提供了新方法。本文将这些进展应用于模型越狱任务:诱导对齐的大模型生成有害响应。我们提出一种对抗推理方法,利用损失信号指导测试时计算,实现了在多种对齐大模型上的当前最优攻击成功率,即使面对那些以牺牲推理时计算为代价来增强对抗鲁棒性的模型也有效。该方法揭示了大模型的新脆弱性,为发展更鲁棒、更可信的人工智能系统奠定了基础。
原文摘要 · Abstract (English)
As large language models (LLMs) are becoming more capable and widespread, the study of their failure cases is becoming increasingly important. Recent advances in standardizing, measuring, and scaling test-time compute suggest new methodologies for optimizing models to achieve high performance on hard tasks. In this paper, we apply these advances to the task of model jailbreaking: eliciting harmful responses from aligned LLMs. We develop an adversarial reasoning approach to automatic jailbreaking that leverages a loss signal to guide the test-time compute, achieving SOTA attack success rates against many aligned LLMs, even those that aim to trade inference-time compute for adversarial robustness. Our approach introduces a new paradigm in understanding LLM vulnerabilities, laying the foundation for the development of more robust and trustworthy AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。