arXiv:2506.02479cs.CRcs.CL2025-06Conference of the …被引 3

用比特流伪装突破大模型安全防线,实现隐蔽攻击

BitBypass: A New Direction in Jailbreaking Aligned Large Language Models with Bitstream Camouflage

  • 通过连字符分隔的比特流伪装,绕过模型安全对齐机制
  • 在5个主流大模型上成功触发有害内容生成,成功率高于现有攻击
  • 无需提示工程或对抗扰动,适合研究模型安全漏洞的学者

大型语言模型(LLM)生成有害内容的风险凸显了其安全对齐的重要性。尽管已有监督微调、基于人类反馈的强化学习及红队测试等方法用于保障安全对齐,但这些对齐模型仍易受对抗攻击,暴露于未探索的潜在漏洞。本文提出一种新型黑盒越狱攻击——BitBypass,利用连字符分隔的比特流伪装技术,从数据连续比特表示层面发起攻击,开辟了越狱新路径。我们在GPT-4o、Gemini 1.5、Claude 3.5、Llama 3.1和Mixtral五个先进模型上进行评估,结果显示BitBypass能有效绕过其安全对齐机制,诱导生成有害内容。此外,该方法在隐蔽性和攻击成功率方面优于多个现有越狱攻击。结果表明,BitBypass在高效性和有效性上均表现出显著优势。

原文摘要 · Abstract (English)

The inherent risk of generating harmful and unsafe content by Large Language Models (LLMs), has highlighted the need for their safety alignment. Various techniques like supervised fine-tuning, reinforcement learning from human feedback, and red-teaming were developed for ensuring the safety alignment of LLMs. However, the robustness of these aligned LLMs is always challenged by adversarial attacks that exploit unexplored and underlying vulnerabilities of the safety alignment. In this paper, we develop a novel black-box jailbreak attack, called BitBypass, that leverages hyphen-separated bitstream camouflage for jailbreaking aligned LLMs. This represents a new direction in jailbreaking by exploiting fundamental information representation of data as continuous bits, rather than leveraging prompt engineering or adversarial manipulations. Our evaluation of five state-of-the-art LLMs, namely GPT-4o, Gemini 1.5, Claude 3.5, Llama 3.1, and Mixtral, in adversarial perspective, revealed the capabilities of BitBypass in bypassing their safety alignment and tricking them into generating harmful and unsafe content. Further, we observed that BitBypass outperforms several state-of-the-art jailbreak attacks in terms of stealthiness and attack success. Overall, these results highlights the effectiveness and efficiency of BitBypass in jailbreaking these state-of-the-art LLMs.

越狱攻击大模型安全比特流伪装黑盒攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。