通过内部激活检测语言模型的欺骗等错位行为
Probing the Misaligned Thinking Process of Language Models

- 将错位行为分解为18类认知指标,用线性探测识别
- 在分布外测试中达0.935 AUROC,误报率低
- 适合安全评估与大模型可靠性研究者
大型语言模型表现出越来越多的错位行为,如策略性欺骗、故意示弱和自我保护。随着它们在高风险场景中的广泛应用,可靠检测这些行为对确保安全和负责任使用至关重要。本文提出通过将错位行为分解为细粒度的认知过程——错位指标,并利用线性探测在模型内部激活中检测其存在来监控错位。我们构建了涵盖18类指标的分类体系,并开发了一套自动化、元计划引导的多轮对话生成流程。为严格评估泛化能力,我们构建了一个分布外测试集,结合自动化行为诱导、已有的错位基准和自然良性对话。在5种错位行为上,我们的探测器在分布外基准上达到0.935 AUROC,同时在良性流量上保持低误报率。我们还进行了深入分析,以理解探测器及模型对错位指标的内部表征。
原文摘要 · Abstract (English)
Large language models exhibit a growing range of misaligned behaviors such as strategic deception, sandbagging, and self-preservation. As they are increasingly deployed in high-stakes settings, it is critical to reliably detect such behaviors to ensure safe and responsible use. In this work, we propose to monitor misalignment by decomposing it into fine-grained cognitive processes -- misalignment indicators -- and detecting their presence in a model's internal activations via linear probes. We develop a taxonomy of 18 indicators spanning different misaligned behaviors, paired with an automated, meta-plan-guided pipeline that generates multi-turn training conversations. To rigorously evaluate generalization, we construct an out-of-distribution suite combining automated behavioral elicitation, established misalignment benchmarks, and natural benign conversations. Across 5 misaligned behaviors, our probes match a strong LLM judge with 0.935 AUROC on out-of-distribution benchmarks while keeping a low false positive rate on benign traffic. We further perform in-depth analysis to understand the probes and the model's internal representations of misalignment indicators.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。