arXiv:2511.16035cs.CLcs.AI2025-11被引 14

构建首个大规模语言模型说谎检测基准,揭示现有技术盲区。

Liars' Bench: Evaluating Lie Detectors for Language Models

  • 设计多维度说谎数据集,覆盖不同说谎动机与信念目标。
  • 在7万余条样本上测试发现,现有方法对隐蔽型说谎识别率不足。
  • 适合研究大模型可信性、安全评估与检测算法的开发者使用。

已有研究提出检测大语言模型说谎的方法,即生成其认为虚假的陈述。然而,这些方法通常在狭窄场景下验证,无法涵盖模型可能产生的多样化谎言。本文提出LIARS' BENCH,一个包含72,863个由四个开源模型在七个数据集上生成的谎言与诚实回答的测试平台。该设置涵盖不同类型的谎言,并在两个维度上变化:模型说谎的动机和所针对的信念对象。在该基准上评估三种黑盒与白盒说谎检测方法,发现现有技术系统性地无法识别某些类型谎言,尤其在仅从文本无法判断是否说谎的情境下表现差。总体而言,LIARS' BENCH 揭示了现有方法的局限性,并为推进说谎检测研究提供了实用测试平台。

原文摘要 · Abstract (English)

Prior work has introduced techniques for detecting when large language models (LLMs) lie, that is, generate statements they believe are false. However, these techniques are typically validated in narrow settings that do not capture the diverse lies LLMs can generate. We introduce LIARS' BENCH, a testbed consisting of 72,863 examples of lies and honest responses generated by four open-weight models across seven datasets. Our settings capture qualitatively different types of lies and vary along two dimensions: the model's reason for lying and the object of belief targeted by the lie. Evaluating three black- and white-box lie detection techniques on LIARS' BENCH, we find that existing techniques systematically fail to identify certain types of lies, especially in settings where it's not possible to determine whether the model lied from the transcript alone. Overall, LIARS' BENCH reveals limitations in prior techniques and provides a practical testbed for guiding progress in lie detection.

模型可信性说谎检测大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。