不同欺骗类型需配不同检测探针,通用探针效果有限。
One Probe Won't Catch Them All: Towards Targeted Deception Detection
- 按欺骗类型定制探测器,比通用探针更有效
- 匹配特定类型可提升AUC达0.108,通用仅+0.032
- 组织应根据威胁模型选择对应探针
线性探针是监控AI系统欺骗行为的有前景方法。以往研究显示,基于对比指令对与简单数据集训练的线性分类器可取得良好效果。然而,这些探针在简单场景中仍存在显著失败,包括虚假相关和对非欺骗性回应产生误报。本文表明,欺骗检测本质上具有异质性:单一通用探针仅带来微弱提升(+0.032 AUC),而事后最优分析显示,当探针与特定欺骗类型匹配时,潜力可大幅提升(+0.108 AUC)。合成验证实验表明,若事先知晓欺骗类型,此上限可被提前实现。研究发现,指令对捕捉的是欺骗意图而非内容模式,解释了提示选择主导探针性能(解释70.6%方差)。因此,建议机构应定义具体威胁模型,并部署相应匹配的探针,而非追求通用欺骗检测器。
原文摘要 · Abstract (English)
Linear probes are a promising approach for monitoring AI systems for deceptive behaviour. Previous work has shown that a linear classifier trained on a contrastive instruction pair and a simple dataset can achieve good performance. However, these probes exhibit notable failures even in straightforward scenarios, including spurious correlations and false positives on non-deceptive responses. In this paper, we demonstrate that deception detection is inherently heterogeneous: while a single universal probe achieves modest improvements (+0.032 AUC), post-hoc oracle analysis reveals substantially higher potential (+0.108 AUC) when probes are matched to specific deception types, and synthetic validation experiments suggest this ceiling is achievable a priori when the deception type is known in advance. Our findings reveal that instruction pairs capture deceptive intent rather than content-specific patterns, explaining why prompt choice dominates probe performance (70.6% of variance). Given this heterogeneity, we conclude that organizations should define their specific threat models and deploy appropriately matched probes rather than seeking a universal deception detector.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。