揭示大模型越狱攻击的内部机制,助力增强安全防御。
NeuroBreak: Unveil Internal Jailbreak Mechanisms in Large Language Models
- 通过分层探测追踪有害表示的演化路径。
- 识别安全相关神经元的倾向与上下文贡献。
- 可视化联合呈现神经元角色与协作关系,支持因果验证。
越狱攻击可绕过大语言模型(LLMs)的安全对齐机制,诱导有害输出,但庞大的参数空间使得诊断其失效机理极为困难。我们提出NeuroBreak,一个视觉分析系统,帮助专家从层级语义逐步剖析越狱机制至神经元层级行为。分层探测管道追踪有害表示在各层的演变过程,双维特征-行为分类揭示每个安全相关神经元的内在倾向与上下文贡献。通过定制化可视化设计实现可解释性:任务驱动的探测投影揭示安全决策边界,双流语义演化流追踪跨层语义转移,特征-行为和弦图将神经元角色、归因分数与协同关系统一于单一视图,并支持原位因果验证。定量评估与案例研究显示,NeuroBreak能有效揭示安全失效原因,并为强化LLM防御提供可操作洞察。
原文摘要 · Abstract (English)
Jailbreak attacks bypass the safety alignment of large language models (LLMs) to elicit harmful outputs, yet the vast parameter space makes diagnosing the underlying failure mechanisms extremely challenging. We present NeuroBreak, a visual analytics system that helps experts progressively unpack jailbreak mechanisms from layer-level semantics down to neuron-level behaviors. A layer-wise probing pipeline traces how harmful representations evolve across layers, while a dual-dimensional character--behavior categorization reveals each safety-related neuron's inherent tendency and contextual contribution. These analyses are made interpretable through tailored visualization designs: a task-driven probing projection that reveals safety decision boundaries, a dual-stream semantic evolution flow that traces cross-layer semantic shifts, and a character--behavior chord graph that unifies neuron roles, attribution scores, and collaborative relations in a single view with in-situ causal verification. Quantitative evaluations and case studies show that NeuroBreak uncovers safety failure causes and provides actionable insights for strengthening LLM defenses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。