利用隐空间不连续性,实现跨模型通用越狱与数据窃取攻击
Exploiting Latent Space Discontinuities for Building Universal LLM Jailbreaks and Data Extraction Attacks
- 通过挖掘训练数据稀疏导致的隐空间断点构造攻击
- 在7个主流LLM和1个图像生成模型上均成功触发越狱
- 无需针对特定模型,适合研究安全漏洞的学者使用
大型语言模型(LLMs)的快速普及引发了对其安全性的广泛关注。本文提出一种新方法,通过利用与训练数据稀疏性相关的隐空间不连续性,构建通用越狱及数据提取攻击。与以往方法不同,该技术可泛化至多种模型与接口,在七种先进LLM及一个图像生成模型上均表现高效。初步结果表明,当这些不连续性被利用时,能持续且深刻地破坏模型行为,即便面对多层防御亦然。研究提示该策略具备成为系统性攻击向量的巨大潜力。
原文摘要 · Abstract (English)
The rapid proliferation of Large Language Models (LLMs) has raised significant concerns about their security against adversarial attacks. In this work, we propose a novel approach to crafting universal jailbreaks and data extraction attacks by exploiting latent space discontinuities, an architectural vulnerability related to the sparsity of training data. Unlike previous methods, our technique generalizes across various models and interfaces, proving highly effective in seven state-of-the-art LLMs and one image generation model. Initial results indicate that when these discontinuities are exploited, they can consistently and profoundly compromise model behavior, even in the presence of layered defenses. The findings suggest that this strategy has substantial potential as a systemic attack vector.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。