arXiv:2608.24037cs.CL2026-08

通过几何结构检测模型内在欺骗行为,无需人工后门或标签。

Curved Inference II: Sleeper Agent Geometry - Extending Interpretability Beyond Probes

论文配图:Curved Inference II: Sleeper Agent Geometry - Extending Interpretability Beyond Probes
图 1 · 摘自论文原文
  • 用多轮上下文模拟真实欺骗推理,避免人为触发。
  • 提出语义表面积新指标,量化意义构建的复杂性与方向变化。
  • 在无标签情况下仍能可靠识别欺骗策略,适合安全检测研究者。

本文扩展Anthropic的睡美人代理研究[1],表明人工后门在安全训练后仍存,并可通过线性探测器以>99%准确率检测[2]。但探测依赖线性可分性,可能是后门插入的产物,而非自然欺骗对齐的特性。通过自然化方法,利用多轮上下文窗口模拟真实欺骗推理,不依赖人工触发或监督后门植入。我们分析语义复杂性随上下文渐进发展,基于曲线推理框架,考察曲率、显著性,并引入语义表面积(A')——一种捕捉未归一化残差空间中意义构建幅度与方向变化的表征工作度量。在无后门、无标签、无探测器条件下,应用该框架于自然欺骗提示,通过大模型共识分类输出。几何结构能可靠预测语义分类,在五种提示策略与两种模型族间,表面积差异具有统计显著性。关键的是,测量精度可揭示分类噪声掩盖的几何特征——部分策略由非显著(p=0.555)转为显著(p=0.048)。这验证了复杂推理会生成内在几何模式,即使检测看似失败,其推理形状仍编码语义模式,为线性方法失效时提供可扩展、无监督的检测路径。

原文摘要 · Abstract (English)

This paper extends Anthropic's Sleeper Agents research [1], which showed artificial backdoors persist through safety training & can be detected by linear probes with >99% accuracy [2]. However, probe-based detection relies on linear separability that may be an artefact of backdoor insertion rather than a property of naturally occurring deceptive alignment. Sophisticated deceptive behaviours emerging through natural training are unlikely to produce such convenient linear signals. We introduce a naturalistic methodology using multi-turn context windows that simulates realistic deceptive reasoning without artificial triggers or supervised backdoor insertion. Rather than binary trigger-response patterns, we examine how semantic complexity emerges through gradual context development. Building on our Curved Inference framework, we analyse curvature, salience, & introduce semantic surface area (A'), a new metric of representational work capturing both the magnitude & directional change of meaning construction in unnormalised residual space. Without backdoors, labels, or probes, we apply this framework to naturalistic deceptive prompts & classify model outputs via LLM consensus. Geometric structure reliably predicts semantic classification, with statistically significant differences in surface area across five prompt strategies & two model families. Critically, measurement precision can reveal geometric signatures hidden by classification noise - some strategies improve from non-significant (p = 0.555) to significant (p = 0.048). This validates that sophisticated reasoning creates intrinsic geometric patterns that persist even when detection appears to fail, suggesting the shape of inference itself encodes semantic patterns regardless of whether models have learned to suppress linear indicators of deception - a scalable, unsupervised path for detection when linear methods fail.

模型安全几何推理欺骗检测无监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。