arXiv:2605.20241cs.LGcs.AI2026-05

通过分层边际几何分析大模型安全提示的判断机制

Geometry-Lite: Interpretable Safety Probing via Layer-Wise Margin Geometry

  • 构建分层边际几何探针,解析每层安全信号的分布特征
  • 发现最终边界位置和异常侧层占用是检测性能关键
  • 适合研究模型安全决策逻辑的开发者与安全研究人员

大语言模型的提示级安全探针利用隐藏状态区分安全与不安全提示,但高平均检测性能未能揭示其分离的几何结构。本文研究这一问题,提出Geometry-Lite:一个紧凑的提示级探针,将各层最终提示标记表示映射至质心、局部邻域和有监督线性边界的带符号边际,并通过边界位置、层间变化及粗略形状总结边际谱。在九个指令微调主干模型(1.2B–70B)和七个安全基准上,Geometry-Lite优于单层探针,且接近原始多层分数堆叠,成为分析多层安全信号的有效工具。结果表明,安全证据主要体现为持续的边界位置几何:最终或极值边际及不安全侧层占用主导整体检测性能。相比之下,有限差分漂移和结构摘要对聚合AUROC贡献较小,尽管漂移可在低误报率阈值下提供微弱召回修正。在基准分布偏移下,优化线性边界在训练混合数据上表现锐利,而类别条件均值几何在预定义困难保留子集上保持分离更稳定。总体而言,提示级安全证据并非主要依赖层间运动信号,而是持久的层内边际几何,其有效成分与读出偏差在决策临界区域中显现。

原文摘要 · Abstract (English)

Prompt-level safety probes for large language models use hidden-state representations to separate safe from unsafe prompts, but strong average detection performance does not explain the geometry of this separation. In particular, it remains unclear how safety evidence is formed across layers, which aspects of that layer-wise geometry support low-false-positive decisions, and which geometric biases remain stable under benchmark shift. We study this as an empirical decomposition problem and introduce Geometry-Lite, a compact prompt-level probe that maps each layer's final prompt-token representation to signed margins under centroid, local-neighborhood, and supervised linear-boundary readouts, then summarizes the resulting margin profiles by boundary position, layer-to-layer change, and coarse shape. Across nine instruction-tuned backbones ($1.2$B--$70$B) and seven safety benchmarks, Geometry-Lite improves over single-layer probes while remaining close to raw multi-layer score stacking, making it a useful instrument for analyzing the multi-layer safety signal. The decomposition shows that safety evidence is expressed primarily through persistent boundary-position geometry: final or extremal margins and unsafe-side layer occupancy dominate aggregate detection performance. In contrast, finite-difference drift and structural summaries add little to pooled AUROC, although drift can provide small recall-oriented corrections under shifted low-FPR thresholds. Under benchmark shift, optimized linear boundaries are sharp on the training mixture, whereas class-conditional mean geometry retains separation more reliably on a predefined hard held-out subset. Overall, prompt-level safety evidence is not primarily a layer-to-layer motion signal, but a persistent layer-wise margin geometry whose useful components and readout-level biases become visible in decision-critical regimes.

安全探针分层分析边际几何大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。