arXiv:2501.07217cs.CL2025-01被引 5

提出首个嵌入式谎言数据集,用大模型识别真假混杂的谎言。

When lies are mostly truthful: automated verbal deception detection for embedded lies

  • 构建2088条含嵌入谎言的语句数据集,真实还原说谎场景。
  • 大模型在真假混杂语句上仅达64%分类准确率,难分真假。
  • 适合研究心理、语言分析与反欺骗技术的学者参考。

背景:传统言语欺骗检测依赖完整真或假陈述,但现实中的谎言常夹杂在真实叙述中。本文首次构建了包含2,088条真/假陈述的数据集,其中每条陈述均包含被标注的嵌入式谎言。采用被试内设计,参与者先描述个人经历的真实陈述,再改写为含嵌入谎言的虚假陈述,并标注谎言的中心性、欺骗程度和来源。结果显示,微调后的Llama-3-8B语言模型在区分真实陈述与含嵌入谎言的陈述时,准确率为64%。个体差异、语言特征与可解释性分析表明,嵌入式谎言之所以难以识别,是因为其与真实陈述高度相似。典型虚假陈述中约2/3内容为真实信息,1/3为嵌入谎言,且多源自过往亲身经历,与真实版本的语言差异极小。本研究提供该数据集作为新资源,推动嵌入式谎言检测的研究发展。

原文摘要 · Abstract (English)

Background: Verbal deception detection research relies on narratives and commonly assumes statements as truthful or deceptive. A more realistic perspective acknowledges that the veracity of statements exists on a continuum with truthful and deceptive parts being embedded within the same statement. However, research on embedded lies has been lagging behind. Methods: We collected a novel dataset of 2,088 truthful and deceptive statements with annotated embedded lies. Using a within-subjects design, participants provided a truthful account of an autobiographical event. They then rewrote their statement in a deceptive manner by including embedded lies, which they highlighted afterwards and judged on lie centrality, deceptiveness, and source. Results: We show that a fined-tuned language model (Llama-3-8B) can classify truthful statements and those containing embedded lies with 64% accuracy. Individual differences, linguistic properties and explainability analysis suggest that the challenge of moving the dial towards embedded lies stems from their resemblance to truthful statements. Typical deceptive statements consisted of 2/3 truthful information and 1/3 embedded lies, largely derived from past personal experiences and with minimal linguistic differences with their truthful counterparts. Conclusion: We present this dataset as a novel resource to address this challenge and foster research on embedded lies in verbal deception detection.

谎言检测嵌入谎言语言模型数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。