研究人类为何容易被大模型伪造信息误导,发现熟悉假新闻的人反而更擅长识别。
Eroding the Truth-Default: A Causal Analysis of Human Susceptibility to Foundation Model Hallucinations and Disinformation in the Wild
- 用因果模型分析人类对大模型幻觉的敏感度,分离真实性和来源判断
- 熟悉假新闻者识别能力更强(相关系数0.35),但政治立场影响极小(-0.10)
- 大模型输出因流畅性绕过人类溯源机制,适合认知安全与信息治理研究者
随着基础模型(FMs)接近人类语言流利度,区分合成内容与真实内容已成为可信网络智能的关键挑战。本文提出JudgeGPT和RogueGPT双轴框架,将‘真实性’与‘来源归属’解耦,以探究人类易受误导的机制。基于对五种模型(含GPT-4、Llama-2)共918次评估,采用结构因果模型(SCMs)构建可检验的因果假设。结果表明,政治倾向与检测表现关联极弱(r = -0.10);而‘假新闻熟悉度’成为候选中介变量(r = 0.35),提示暴露经历可能充当人类判别器的对抗性训练。我们发现‘流畅性陷阱’:GPT-4输出(HumanMachineScore: 0.20)可绕过源监控机制,使其与真人文本无法区分。研究建议‘预打假’干预应聚焦认知源监控,而非基于人口统计特征的分组。
原文摘要 · Abstract (English)
As foundation models (FMs) approach human-level fluency, distinguishing synthetic from organic content has become a key challenge for Trustworthy Web Intelligence. This paper presents JudgeGPT and RogueGPT, a dual-axis framework that decouples "authenticity" from "attribution" to investigate the mechanisms of human susceptibility. Analyzing 918 evaluations across five FMs (including GPT-4 and Llama-2), we employ Structural Causal Models (SCMs) as a principal framework for formulating testable causal hypotheses about detection accuracy. Contrary to partisan narratives, we find that political orientation shows a negligible association with detection performance ($r=-0.10$). Instead, "fake news familiarity" emerges as a candidate mediator ($r=0.35$), suggesting that exposure may function as adversarial training for human discriminators. We identify a "fluency trap" where GPT-4 outputs (HumanMachineScore: 0.20) bypass Source Monitoring mechanisms, rendering them indistinguishable from human text. These findings suggest that "pre-bunking" interventions should target cognitive source monitoring rather than demographic segmentation to ensure trustworthy information ecosystems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。