用欺骗理论指导AI判断谎言,发现效果有限但偏差可调。
Theory-Guided Deception Detection: A RAG-Based Artificial Intelligence Exploration
- 基于欺骗理论构建RAG模型,对比基准模型的判断差异。
- 准确率约54.6%,与人类水平相当,理论影响小但偏见差异大。
- 不同理论导致偏差从32%到88%,适合研究认知偏见的学者。
本研究基于主流欺骗理论构建了七个检索增强生成(RAG)模型,对比其与基准模型在欺骗判断上的表现。在来自五个公开数据集的700条陈述上,结合四个大语言模型(gpt-4o、claude-sonnet-4-6、ollama/llama3、deepseek-v4-flash)及两种运行方式(RAG vs. baseline),共生成39,200次判断。结果显示,RAG模型准确率为54.5%,基准模型为54.6%,二者无统计差异;但RAG模型(57.0%)比基准模型(59.7%)更少倾向于相信真相,效应量较小。理论视角对准确率影响不大,但显著影响反应偏差:从验证性方法的32.2%高度偏向谎言,到真相默认理论的88.1%高度偏向真实。内容因素与模型差异也进一步调节结果。当前参数下,理论引导的AI判断不可靠,但未来通过更多数据、模型测试和理论-数据匹配或具潜力。
原文摘要 · Abstract (English)
The current work developed seven Retrieval-Augmented Generation (RAG) models based on leading deception theories and compared how deception judgments were made relative to baseline models. Across 700 statements drawn from five published deception datasets, four large language models (gpt-4o, claude-sonnet-4-6, ollama/llama3, deepseek-v4-flash), and two run-types (RAG vs. baseline), a total of 39,200 deception judgments were rendered. Detection accuracies were consistent with typical human accuracies and not statistically different across RAG (54.5%) and baseline models (54.6%). RAG-based models (57.0%) were less truth-biased than baseline models (59.7%), but the effect size was quite small. Theoretical perspective mattered little for accuracy yet mattered substantially for response bias, which ranged from highly lie-biased (the verifiability approach, 32.2%) to highly truth-biased (truth-default theory, 88.1%). Content effects and model effects further moderated the results. Theory-guided AI judgments are unreliable with current parameters, yet they might show promise with additional datasets, model testing, and theory-to-data matching.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。