arXiv:2601.22871cs.CYcs.AI2026-01中稿 · ACM TheWebConf '26…被引 7

研究人类为何容易被大模型伪造信息误导,发现熟悉假新闻的人反而更擅长识别。

Eroding the Truth-Default: A Causal Analysis of Human Susceptibility to Foundation Model Hallucinations and Disinformation in the Wild

  • 用因果模型分析人类对大模型幻觉的敏感度,分离真实性和来源判断
  • 熟悉假新闻者识别能力更强(相关系数0.35),但政治立场影响极小(-0.10)
  • 大模型输出因流畅性绕过人类溯源机制,适合认知安全与信息治理研究者

随着基础模型(FMs)接近人类语言流利度,区分合成内容与真实内容已成为可信网络智能的关键挑战。本文提出JudgeGPT和RogueGPT双轴框架,将‘真实性’与‘来源归属’解耦,以探究人类易受误导的机制。基于对五种模型(含GPT-4、Llama-2)共918次评估,采用结构因果模型(SCMs)构建可检验的因果假设。结果表明,政治倾向与检测表现关联极弱(r = -0.10);而‘假新闻熟悉度’成为候选中介变量(r = 0.35),提示暴露经历可能充当人类判别器的对抗性训练。我们发现‘流畅性陷阱’:GPT-4输出(HumanMachineScore: 0.20)可绕过源监控机制,使其与真人文本无法区分。研究建议‘预打假’干预应聚焦认知源监控,而非基于人口统计特征的分组。

原文摘要 · Abstract (English)

As foundation models (FMs) approach human-level fluency, distinguishing synthetic from organic content has become a key challenge for Trustworthy Web Intelligence. This paper presents JudgeGPT and RogueGPT, a dual-axis framework that decouples "authenticity" from "attribution" to investigate the mechanisms of human susceptibility. Analyzing 918 evaluations across five FMs (including GPT-4 and Llama-2), we employ Structural Causal Models (SCMs) as a principal framework for formulating testable causal hypotheses about detection accuracy. Contrary to partisan narratives, we find that political orientation shows a negligible association with detection performance ($r=-0.10$). Instead, "fake news familiarity" emerges as a candidate mediator ($r=0.35$), suggesting that exposure may function as adversarial training for human discriminators. We identify a "fluency trap" where GPT-4 outputs (HumanMachineScore: 0.20) bypass Source Monitoring mechanisms, rendering them indistinguishable from human text. These findings suggest that "pre-bunking" interventions should target cognitive source monitoring rather than demographic segmentation to ensure trustworthy information ecosystems.

大模型幻觉认知机制信息可信度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。