人类无法可靠区分大模型生成的新闻真伪,检测失效。
Can Humans Tell? A Dual-Axis Study of Human Perception of LLM-Generated News
- 用双轴评估平台测量人类对文本来源和真实性的判断
- 2318次判断显示人类识别率不高于随机水平(p>0.05)
- 专家经验提升判断准确率,但连续评估后会因疲劳下降
我们通过JudgeGPT平台,独立测量用户对新闻来源(人写/模型写)和真实性(真实/虚假)的判断。基于1054名参与者对6种大语言模型生成内容的2318次评估,发现:(1)人类无法可靠区分机器生成与人工撰写文本(p > .05,Welch's t检验);(2)该现象在所有测试模型中均存在,包括参数仅为7B的开源模型;(3)自我报告的专业领域经验能预测判断准确性(r = .35, p < .001),而政治倾向无显著影响(r = -.10, n.s.);(4)聚类分析揭示出两类响应策略(“怀疑派”与“相信派”);(5)约30轮连续评估后,准确率因认知疲劳下降。结论是:人类无法可靠识别。这表明用户端检测不可行,需系统级应对,如加密内容溯源。
原文摘要 · Abstract (English)
Can humans tell whether a news article was written by a person or a large language model (LLM)? We investigate this question using JudgeGPT, a study platform that independently measures source attribution (human vs. machine) and authenticity judgment (legitimate vs. fake) on continuous scales. From 2,318 judgments collected from 1,054 participants across content generated by six LLMs, we report five findings: (1) participants cannot reliably distinguish machine-generated from human-written text (p > .05, Welch's t-test); (2) this inability holds across all tested models, including open-weight models with as few as 7B parameters; (3) self-reported domain expertise predicts judgment accuracy (r = .35, p < .001) whereas political orientation does not (r = -.10, n.s.); (4) clustering reveals distinct response strategies ("Skeptics" vs. "Believers"); and (5) accuracy degrades after approximately 30 sequential evaluations due to cognitive fatigue. The answer, in short, is no: humans cannot reliably tell. These results indicate that user-side detection is not a viable defense and motivate system-level countermeasures such as cryptographic content provenance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。