测试大模型欺骗检测探针的鲁棒性,发现其性能随风格变化严重下降。
Pressure-Testing Deception Probes in LLMs: Scaling, Robustness, and the Geometry of Deceptive Representations

- 通过跨域转移和多维探测分析,验证欺骗信号非单一方向
- 清洁数据上探测器AUROC超0.998,风格迁移后降至0.45以下
- 风格增强后恢复高精度,说明问题出在训练分布而非模型规模
线性探针被广泛用于检测大语言模型中的欺骗行为,但在干净基准上报告的AUROC超过0.96,而在分布外数据上却迅速失效。本文系统地对Gemma 3系列(1B-27B参数)模型的探针检测能力进行压力测试,诊断其失败原因而非仅描述现象。检验四种欺骗编码假设:(1) 单一线性方向,(2) 多维子空间,(3) 凸锥包,(4) 熵代理。实验设计包括跨领域转移矩阵、多维探针分析与置换零假设、熵残差测试及8种风格扰动下的干扰评估。结果表明:(a) 探针在清洁数据上达到≥0.998的AUROC,但在风格迁移下崩溃;风格增强探针在未见风格上恢复平均0.979-0.983的检测性能;(b) 单方向假设被否定(k=1时仅捕捉0.61-0.80 AUROC),跨域失败源于几何特性而非层不匹配;(c) 熵代理假设被否定(最大|rho|=0.454,残差化后Δ-AUROC最大为0.004);(d) 欺骗信号不构成显著线性子空间(每领域k*=0),但多维探针(k≥5)可通过分布式亚阈值特征恢复信号。探针脆弱性源于分布狭窄,而非架构限制:在4B和27B模型上,风格增强探针均能恢复近完美检测,证明反向缩放现象是训练分布的产物,而非真实的尺度依赖效应。
原文摘要 · Abstract (English)
Linear probes trained on LLM activations are increasingly proposed as deception-detection metrics, yet report AUROC exceeding 0.96 on clean benchmarks while collapsing under distributional shift. This paper systematically pressure-tests probe-based metrics across the Gemma 3 model family (1B-27B parameters), diagnosing why they fail rather than merely documenting that they fail. We test four hypotheses about deception encoding: (1) single linear direction, (2) multi-dimensional subspace, (3) convex conic hull, (4) entropy proxy. Our design includes cross-domain transfer matrices, multi-dimensional probe analysis with permutation null baselines, entropy-residualization tests, and distractor evaluations across 8 stylistic shifts. We find that: (a) probes achieve near-perfect AUROC (>=0.998) on clean data but collapse under stylistic shifts; style-augmented probes recover near-perfect detection (mean AUROC 0.979-0.983) on unseen styles; (b) the single-direction hypothesis is rejected (k=1 captures only 0.61-0.80 AUROC), with cross-domain transfer failure confirmed as geometric rather than layer-mismatch-driven; (c) the entropy-proxy hypothesis is rejected (max |rho|=0.454, max Delta-AUROC after residualization=0.004); and (d) deception does not form a significant linear subspace (per-domain k*=0), yet multi-dimensional probes (k>=5) recover the signal through distributed sub-threshold features. Probe fragility reflects distributional narrowness rather than an architectural limitation: style-augmented probes recover near-perfect detection at both 4B and 27B, establishing that the inverse scaling pattern is a training-distribution artifact rather than a genuine scale-dependent phenomenon.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。