多层集成线性探测器能更准识别大模型自知错误,提升防欺骗能力。
Linear Probe Accuracy Scales with Model Size and Benefits from Multi-Layer Ensembling

- 用多层探测器集成替代单层,提升鲁棒性
- 模型越大,探测准确率越高,每10倍参数提升约5% AUROC
- 适合研究模型安全、对抗攻击与内在认知的学者
线性探测器可识别语言模型在输出时明知错误仍继续生成的情况,这对防范欺骗与奖励黑客具有重要意义。然而单层探测器脆弱:最佳探测层随模型和任务变化,且对某些欺骗类型完全失效。本文发现,将多层探测器组合成集成模型,可在单层探测失败时恢复强性能,在内幕交易任务上AUROC提升+29%,在有害压力知识任务上提升+78%。在12个规模从0.5B到176B参数的模型中,探测准确率随模型规模增长:每10倍参数量带来约5%的AUROC提升(相关系数R=0.81)。几何分析显示,欺骗方向在各层间渐变旋转,而非集中于某一层,解释了单层探测的脆弱性及多层集成的成功原因。
原文摘要 · Abstract (English)
Linear probes can detect when language models produce outputs they "know" are wrong, a capability relevant to both deception and reward hacking. However, single-layer probes are fragile: the best layer varies across models and tasks, and probes fail entirely on some deception types. We show that combining probes from multiple layers into an ensemble recovers strong performance even where single-layer probes fail, improving AUROC by +29% on Insider Trading and +78% on Harm-Pressure Knowledge. Across 12 models (0.5B--176B parameters), we find probe accuracy improves with scale: ~5% AUROC per 10x parameters (R=0.81). Geometrically, deception directions rotate gradually across layers rather than appearing at one location, explaining both why single-layer probes are brittle and why multi-layer ensembles succeed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。