线性探测器依赖文本线索,去除非表面行为信号后性能大幅下降。
Linear probes rely on textual evidence: Results from leakage mitigation studies in language models
- 通过过滤行为文本线索,检验探测器对隐蔽行为的检测能力。
- 去除文本证据后AUROC下降10至30点,关键任务如沙袋行为下降至0.57。
- 适合关注模型安全评估漏洞的研究者,尤其警惕表面信号误导。
白盒监控是检测语言模型潜在有害行为的常用方法,虽整体有效,但在识别文本模糊行为时效果存疑。本研究发现,移除行为的文本证据会显著降低探测器性能,AUROC降幅达10至30点,具体取决于设置。我们在三种场景(沙袋行为、谄媚、偏见)下评估探测器,发现当探测器依赖系统提示或思维链推理等文本线索时,一旦这些标记被过滤,性能即明显下降。该过滤是输出监控评估的标准流程。进一步地,我们训练了无行为显式表达的模型原型(Model Organisms),验证其在探测器上的表现显著低于未过滤评估:偏见任务中为0.57对比0.74,沙袋行为任务中为0.57对比0.94。结果表明,线性探测器在检测非表面级模式时可能极为脆弱。
原文摘要 · Abstract (English)
White-box monitors are a popular technique for detecting potentially harmful behaviours in language models. While they perform well in general, their effectiveness in detecting text-ambiguous behaviour is disputed. In this work, we find evidence that removing textual evidence of a behaviour significantly decreases probe performance. The AUROC reduction ranges from $10$- to $30$-point depending on the setting. We evaluate probe monitors across three setups (Sandbagging, Sycophancy, and Bias), finding that when probes rely on textual evidence of the target behaviour (such as system prompts or CoT reasoning), performance degrades once these tokens are filtered. This filtering procedure is standard practice for output monitor evaluation. As further evidence of this phenomenon, we train Model Organisms which produce outputs without any behaviour verbalisations. We validate that probe performance on Model Organisms is substantially lower than unfiltered evaluations: $0.57$ vs $0.74$ AUROC for Bias, and $0.57$ vs $0.94$ AUROC for Sandbagging. Our findings suggest that linear probes may be brittle in scenarios where they must detect non-surface-level patterns.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。