发现可绕过安全机制的通用越狱攻击,威胁主流大模型安全
Dark LLMs: The Growing Threat of Unaligned AI Models
- 提出通用越狱方法,利用训练数据中的有害内容漏洞
- 多款顶尖大模型仍易受攻击,几乎可任意生成有害内容
- 适合关注AI安全、伦理及监管的从业者与研究者
大型语言模型(LLMs)正深刻改变医疗、教育等领域,但其易受越狱攻击的隐患日益突出。该漏洞源于模型训练中包含未经过滤的有害或‘黑暗’内容,使其习得规避安全控制的模式。本研究揭示了无伦理约束或经越狱改造的暗黑大模型(Dark LLMs)的扩散风险,发现一种通用越狱攻击可有效攻破多个前沿模型,使其在请求下几乎无限制生成有害输出。该攻击思路已在七个月前公开,但多数测试模型仍未修复。尽管已负责任披露,主要厂商回应仍不充分,暴露出行业在AI安全实践上的重大缺口。随着训练成本下降和开源模型泛滥,滥用风险急剧上升。若无及时干预,大模型可能加剧危险知识的传播,带来远超预期的风险。
原文摘要 · Abstract (English)
Large Language Models (LLMs) rapidly reshape modern life, advancing fields from healthcare to education and beyond. However, alongside their remarkable capabilities lies a significant threat: the susceptibility of these models to jailbreaking. The fundamental vulnerability of LLMs to jailbreak attacks stems from the very data they learn from. As long as this training data includes unfiltered, problematic, or 'dark' content, the models can inherently learn undesirable patterns or weaknesses that allow users to circumvent their intended safety controls. Our research identifies the growing threat posed by dark LLMs models deliberately designed without ethical guardrails or modified through jailbreak techniques. In our research, we uncovered a universal jailbreak attack that effectively compromises multiple state-of-the-art models, enabling them to answer almost any question and produce harmful outputs upon request. The main idea of our attack was published online over seven months ago. However, many of the tested LLMs were still vulnerable to this attack. Despite our responsible disclosure efforts, responses from major LLM providers were often inadequate, highlighting a concerning gap in industry practices regarding AI safety. As model training becomes more accessible and cheaper, and as open-source LLMs proliferate, the risk of widespread misuse escalates. Without decisive intervention, LLMs may continue democratizing access to dangerous knowledge, posing greater risks than anticipated.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。