arXiv:2507.12691cs.AIcs.LG2025-07被引 15

通过黑盒到白盒性能提升,评估欺骗探测器的有效性

Benchmarking Deception Probes via Black-to-White Performance Boosts

  • 用白盒与黑盒监控对比,衡量探测器实际效能
  • 现有探测器在黑盒下表现弱但仍有可观察的性能提升
  • 适合关注AI诚实性检测与对抗防御的研究者

AI助手有时会欺骗用户。近期研究使用线性分类器(称为“欺骗探测器”)来区分语言模型在欺骗与诚实回答时的内部激活状态。然而,这些探测器在实际中是否有效,以及能否抵抗欺骗助手的简单规避策略,尚不明确。本文通过比较白盒监控(可访问逐标记探测激活)与黑盒监控(无此访问)的性能差异,以“黑盒到白盒性能提升”为基准来评测欺骗探测器。结果发现,现有探测器虽仅表现出微弱但令人鼓舞的性能提升,表明其具备一定检测能力。

原文摘要 · Abstract (English)

AI assistants will occasionally respond deceptively to user queries. Recently, linear classifiers (called "deception probes") have been trained to distinguish the internal activations of a language model during deceptive versus honest responses. However, it's unclear how effective these probes are at detecting deception in practice, nor whether such probes are resistant to simple counter strategies from a deceptive assistant who wishes to evade detection. In this paper, we compare white-box monitoring (where the monitor has access to token-level probe activations) to black-box monitoring (without such access). We benchmark deception probes by the extent to which the white box monitor outperforms the black-box monitor, i.e. the black-to-white performance boost. We find weak but encouraging black-to-white performance boosts from existing deception probes.

欺骗检测模型监控白盒分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。