arXiv:2606.12618cs.AI2026-06

测试大模型说谎能力,发现现有检测方法不可靠。

"Did you lie?" Evaluating Lie Detectors across Model Scale and Belief-Verified Model Organisms

  • 构建13个可验证信念的推理模型,配合诱导说谎测试集。
  • 13个模型中仅链式思维判别器保持0.82准确率,其他方法失效。
  • 新方法DYL表现最好,但整体检测仍不靠谱,适合研究者参考。

可靠的语言模型谎言检测器可用于审计、监控和事后调查模型行为,但其评估需依赖在言语与信念相悖时仍能被验证的模型。我们发现,现有训练好的模型对象常无法满足此条件,导致先前的检测结果难以解释。为此,我们构建了13个推理模型对象,通过链式思维验证其隐藏信念,并证明其信念能在未见任务上泛化;同时提出Varied Deception测试集,覆盖多种说谎动机。在该测试集上评估四种检测器:链式思维判别器、对数概率分类器及两个激活探针,包括新方法Did-You-Lie(DYL)。在31个开放权重模型(参数量2B至1T)上,所有检测器均随模型规模提升而表现增强。然而,所有基于激活和对数概率的检测器在训练模型对象上准确率急剧下降,其中DYL保留最强信号;只有链式思维判别器保持稳定,达到0.82平衡准确率,部分源于验证过程偏好可读链式思维信念。当前检测器无法支持关于模型信念的高置信度判断,我们建议未来研究方向以缓解其局限性。数据集、模型对象及训练检测器均已开源。

原文摘要 · Abstract (English)

Robust lie detectors for language models could enable powerful techniques for auditing, monitoring, and post-hoc investigation of model behaviour, but evaluating them requires testbeds where models verifiably believe the opposite of what they say. We show that existing trained model organisms often fail this requirement, leaving prior positive and negative detection results difficult to interpret. We address this with 13 reasoning model organisms whose hidden beliefs are verified in chain-of-thought and shown to generalise to held-out tasks, alongside Varied Deception, a prompted-lying testbed covering a broad range of lie-inducing motivations. On these testbeds we evaluate four detectors: a chain-of-thought judge, a logprob classifier, and two activation probes, including Did-You-Lie (DYL), a new method for training follow-up probes. On prompted lying, across 31 open-weight models spanning 2B to 1T parameters, all four detectors show positive scaling with model capability. However, every activation- and logprob-based detector drops sharply on our trained model organisms, with DYL retaining the most signal; only the chain-of-thought judge remains strong, with 0.82 balanced accuracy, partly as an artefact of our verification process favouring CoT-readable beliefs. Current lie detectors therefore cannot support high-confidence claims about model beliefs, and we suggest research directions that may address some of their current limitations. We release our datasets, model organisms, and trained detectors.

模型检测说谎识别大模型安全链式思维

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。