AI写作泛滥引发读者反扑,误判成风却成社交圈层壁垒。
"That's AI Slop, You Bot!" Studying Accusations, Evidence, and Credibility in Online Discourse Towards LLM-Generated Comments
- 用大数据分析2500万条评论,发现'AI垃圾'标签激增十倍
- 实证显示:人类文本被误判为AI,因特征差异不决定指控
- 指控本质是身份认同的排他性信号,非技术检测手段可解
生成式AI使流畅文字廉价可得,打破了好文等于深思的传统承诺。我们分析了2023至2026年间来自Hacker News和Reddit的2500万条评论,结合对7500条样本指控的LLM判断、情感轨迹分析、300条确认指控的语用行为编码,以及指控者与非指控者母评论的匹配对照实验。发现两平台贬义标签中‘AI垃圾’占比上升逾十倍,而2022年前用于指代虚假言论的术语(如shill、astroturf)未见增长。这一转变反映出将可疑或看似不真实文本统称为‘AI垃圾’的迅速蔓延趋势,该表述已占贬义提及的94%。评论语气从讽刺转向排他性门禁与结构性抗议。关键发现来自对照测试:虽有统计特征能区分AI与人类文本,但这些特征无法预测哪些人类文本会被指控为AI。新指控实质上是社会性门禁机制,而非真实检测。本研究拓展了信号理论,揭示即使不准确,替代性社会信号仍可能在非专家层面持续演化,因底层检测难题难以破解。这表明读者端对AI的影响,与写作者端截然不同。检测技术无法解决此动态,因指控的核心功能正从识别生成内容转向群体身份表达。
原文摘要 · Abstract (English)
Generative AI has made fluent prose cheap to produce, breaking the old promise to readers that good writing meant real thinking. How have readers responded, and what can this tell us about changing anti-AI attitudes? We analyzed 25 million comments from Hacker News and Reddit (2023-2026), combining LLM judgment on 7,500 sampled accusations of AI use, sentiment trajectories, speech-act coding of 300 confirmed accusations of AI use, and a matched-control test of accused versus non-accused parent comments. We found that the pejorative-label share of accusations rose more than tenfold on both platforms while a placebo vocabulary of pre-2022 inauthenticity terms (shill, astroturf) did not. This shift reflected a fast-growing trend of branding any suspicious or seemingly inauthentic prose as "AI slop". The slop frame now constitutes 94 percent of pejorative mentions, with the dominant comments shifting in tone from mockery toward gatekeeping and structural protest. The key surprise comes from a matched-control test which found that prose features that statistically distinguish AI from human text do not predict which human text gets accused as AI. The new accusations work as social gatekeeping of perceived authenticity without actually screening for AI. This research extends signaling theory by showing that substitute signals used socially can grow even when inaccurate if the underlying detection problem cannot be solved at the non-expert level. It shows that AI's effects on writing from the reader side are distinct from those on the production (writer) side. Detection technology cannot resolve this dynamic because the social function of accusations is increasingly to perform social gatekeeping and in-group signaling as opposed to identifying AI-generated writing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。