研究发现,检测大模型谎言效果受谎言类型和数据选择影响极大。
Beyond Liars' Bench: The Impact of Lie Typology, Depth, and Sparsity on Deception Detection in LLMs
- 用多种谎言类型数据训练探测器,提升泛化能力
- 深层表征不总是更好,最优深度依赖数据集
- 适合关注大模型安全与虚假内容检测的研究者
训练探测器以识别大语言模型的欺骗性输出仍是开放问题。现有研究表明,探测器在跨领域场景下表现不佳——在一种谎言类型上训练的效果难以迁移至其他类型。本文系统研究了表征深度、探测器表达能力、稀疏特征表示以及训练数据中的谎言类型对检测性能的影响。为此,我们在标准基准数据外引入包含虚构、隐瞒、夸大等多种谎言类型的补充数据集。通过七种探测器的实验表明:最优表征深度高度依赖数据集;更复杂的探测器仅在部分情况下优于线性基线;稀疏自编码器特征与密集隐藏状态表现相当。最终发现,训练数据及谎言类型的选择显著影响可检测性,揭示欺骗检测是高度依赖表征的问题。
原文摘要 · Abstract (English)
Training probes to detect deceptive outputs from large language models is still an open problem. Recent work has demonstrated that detection probes fail especially in out-of-domain scenarios -- training on one type of lie does not transfer well to deception scenarios involving other types of lies. In this work, we conduct a systematic study on how various factors impact detection performance: representation depth, probe expressivity, sparse feature representations, and the lie typology of the training data. To this end, we augment standard benchmark training data with a supplementary dataset containing diverse types of deception, including fabrication, omission, and exaggeration examples. Analyzing these factors across seven probe types, our experimental results show that the optimal representation depth is highly dataset-dependent, more expressive probes provide only selective gains over linear baselines, and sparse autoencoder features perform similarly to dense hidden states. Ultimately, we demonstrate that the choice of training data and lie typology substantially changes detectability, highlighting that deception detection is a highly representation-dependent problem.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。