用社会科学研究方法,检测大模型生成政治言论的不真实感。
The Algorithmic Caricature: Auditing LLM-Generated Political Discourse Across Crisis Events

- 对比真实与生成言论在情绪、结构、词汇上的群体差异
- 发现生成内容更负面、更刻板,情感分布更单一
- 提出‘讽刺差距’指标,适合评估生成内容的社会真实性
大型语言模型(LLMs)可大规模生成流畅的政治文本,引发危机期间合成话语的担忧。现有检测多依赖句子级特征如困惑度、突发性或标记异常,但随着生成系统进步,这些信号可能减弱。本文采用计算社会科学视角,考察合成政治话语是否像真实在线群体行为。构建了涵盖九个危机事件(包括新冠疫情、1月6日国会山事件、2020与2024年美国大选、多布斯案/Roe诉韦德案、2020年黑人命也是命抗议、美国中期选举、犹他州枪击案及美伊战争)的配对语料库,共1,789,406条帖子。比较真实平台数据与同情境下生成内容在情感强度、结构规律性、词汇-意识形态框架及跨事件依赖性四个维度的表现。结果显示:合成话语虽流畅,但在群体层面不真实——整体更负面、情感分布更集中、结构更规则、词汇更抽象;而真实话语则表现出更广的情绪波动、更长尾的结构分布和更多上下文相关的口语化表达。差异程度随事件类型变化:在快速传播、去中心化的危机中更显著,而在正式或制度化事件中较弱。通过‘讽刺差距’(Caricature Gap)量化该现象。研究指出,合成政治话语的主要缺陷不在语法或流畅性,而在于缺乏群体现实性。群体级审计补充传统检测方法,并为评估生成话语的社会真实性提供计算社会科学框架。
原文摘要 · Abstract (English)
Large Language Models (LLMs) can generate fluent political text at scale, raising concerns about synthetic discourse during crises and social conflict. Existing AI-text detection often focuses on sentence-level cues such as perplexity, burstiness, or token irregularities, but these signals may weaken as generative systems improve. We instead adopt a Computational Social Science perspective and ask whether synthetic political discourse behaves like an observed online population. We construct a paired corpus of 1,789,406 posts across nine crisis events: COVID-19, the Jan. 6 Capitol attack, the 2020 and 2024 U.S. elections, Dobbs/Roe v. Wade, the 2020 BLM protests, U.S. midterms, the Utah shooting, and the U.S.-Iran war. For each event, we compare observed discourse from social platforms with synthetic discourse generated for the same context. We evaluate four dimensions: emotional intensity, structural regularity, lexical-ideological framing, and cross-event dependency, using mean gaps and dispersion evidence. Across events, synthetic discourse is fluent but population-level unrealistic. It is generally more negative and less dispersed in sentiment, structurally more regular, and lexically more abstract than observed discourse. Observed discourse instead shows broader emotional variation, longer-tailed structural distributions, and more context-specific, colloquial lexical markers. These differences are event-dependent: larger for fast-moving, decentralized crises and smaller for formal or institutionally mediated events. We summarize them with a simple event-level measure, the Caricature Gap. Our findings suggest that the main limitation of synthetic political discourse is not grammar or fluency, but reduced population realism. Population-level auditing complements traditional text-detection and provides a CSS framework for evaluating the social realism of generated discourse.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。