arXiv:2510.06594cs.CL2025-10被引 3

分析大模型内部层响应,找逃逸攻击的识别线索。

Do Internal Layers of LLMs Reveal Patterns for Jailbreak Detection?

  • 考察模型各隐藏层对恶意与正常提示的响应差异。
  • 发现GPT-J和Mamba2在不同层呈现明显不同的行为模式。
  • 为构建更鲁棒的检测系统提供新思路,适合安全研究者。

随着对话式大语言模型的普及,模型逃逸攻击已成为紧迫问题。攻击者通过精心设计的提示诱导模型输出受限或敏感内容,这种策略被称为逃逸。尽管已有多种防御机制,但攻击者不断演化新方法,现有模型仍无法完全抵御。本研究通过分析大语言模型的内部表征,重点考察隐藏层对逃逸与良性提示的响应差异。以开源模型GPT-J和状态空间模型Mamba2为例,初步揭示了分层行为特征。结果表明,利用模型内部动态有望为逃逸检测与防御提供新方向。

原文摘要 · Abstract (English)

Jailbreaking large language models (LLMs) has emerged as a pressing concern with the increasing prevalence and accessibility of conversational LLMs. Adversarial users often exploit these models through carefully engineered prompts to elicit restricted or sensitive outputs, a strategy widely referred to as jailbreaking. While numerous defense mechanisms have been proposed, attackers continuously develop novel prompting techniques, and no existing model can be considered fully resistant. In this study, we investigate the jailbreak phenomenon by examining the internal representations of LLMs, with a focus on how hidden layers respond to jailbreak versus benign prompts. Specifically, we analyze the open-source LLM GPT-J and the state-space model Mamba2, presenting preliminary findings that highlight distinct layer-wise behaviors. Our results suggest promising directions for further research on leveraging internal model dynamics for robust jailbreak detection and defense.

模型安全逃逸检测内部表征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。