arXiv:2602.19396cs.AI2026-02

通过解耦语义与表述,精准识别隐藏恶意意图的越狱提示。

Hiding in Plain Text: Detecting Concealed Jailbreaks via Activation Disentanglement

  • 在推理时解耦模型激活中的目标与表述因子,分离潜在恶意意图。
  • 新检测器FrameShield在多类大模型上实现高精度越狱攻击识别。
  • 可作为安全防护与模型可解释性研究的基础工具,适合安全团队使用。

大型语言模型仍易受流畅且语义连贯的越狱提示攻击,这类攻击难以通过传统启发式方法检测。攻击者常通过调整请求表述来隐藏真实恶意目标,导致依赖结构特征或目标特定签名的防御失效。为此,我们提出一种自监督框架,在推理阶段解耦大模型激活中的语义因子对。以目标与表述为例,构建了包含可控目标与表述变化的基准数据集GoalFrameBench,用于训练冻结模型下的表示解耦模块ReDAct,提取解耦后的表征。进一步提出基于表述表征的异常检测器FrameShield,实现跨多种大模型家族的模型无关检测,计算开销极低。理论保证与大量实验证明,解耦有效支撑了检测性能。最后,利用解耦作为可解释性探针,揭示目标与表述信号的独立特征分布,确立语义解耦在模型安全与机制可解释性中的基础作用。

原文摘要 · Abstract (English)

Large language models (LLMs) remain vulnerable to jailbreak prompts that are fluent and semantically coherent, and therefore difficult to detect with standard heuristics. A particularly challenging failure mode occurs when an attacker tries to hide the malicious goal of their request by manipulating its framing to induce compliance. Because these attacks maintain malicious intent through a flexible presentation, defenses that rely on structural artifacts or goal-specific signatures can fail. Motivated by this, we introduce a self-supervised framework for disentangling semantic factor pairs in LLM activations at inference. We instantiate the framework for goal and framing and construct GoalFrameBench, a corpus of prompts with controlled goal and framing variations, which we use to train Representation Disentanglement on Activations (ReDAct) module to extract disentangled representations in a frozen LLM. We then propose FrameShield, an anomaly detector operating on the framing representations, which improves model-agnostic detection across multiple LLM families with minimal computational overhead. Theoretical guarantees for ReDAct and extensive empirical validations show that its disentanglement effectively powers FrameShield. Finally, we use disentanglement as an interpretability probe, revealing distinct profiles for goal and framing signals and positioning semantic disentanglement as a building block for both LLM safety and mechanistic interpretability.

越狱检测语义解耦模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。