arXiv:2508.20468cs.CL2025-08Transactions of th…被引 1

构建首个阴谋论认知特征数据集,评估大模型对阴谋话语的脆弱性。

ConspirED: A Dataset for Cognitive Traits of Conspiracy Theories and Large Language Model Safety

  • 基于认知框架标注阴谋论文本的多句片段,提取核心思维特征。
  • 发现大模型易被阴谋论话术误导,即使能识别事实错误仍会模仿其表达。
  • 适合研究虚假信息防御、模型安全与认知偏见的学者使用。

阴谋论削弱公众对科学和机构的信任,且在面对反驳时仍能持续演化并吸收反证。随着人工智能生成的虚假信息日益复杂,理解阴谋论内容中的修辞模式对于制定针对性预辟谣策略及评估人工智能漏洞至关重要。本文提出CONSPIRED(CONSPIR Evaluation Dataset),通过使用CONSPIR认知框架,对来自在线阴谋论文章的多句片段(80-120词)进行标注,捕捉其认知特征。CONSPIRED是首个对阴谋论内容进行通用认知特征标注的数据集。利用该数据集,我们(i)开发了可识别阴谋论特征及主导特征的计算模型;(ii)评估大语言模型对阴谋论输入的鲁棒性。结果表明,大语言模型极易受阴谋论语境影响,即便成功驳斥类似事实核查过的虚假信息,仍会复现其修辞模式。

原文摘要 · Abstract (English)

Conspiracy theories erode public trust in science and institutions while resisting debunking by evolving and absorbing counter-evidence. As AI-generated misinformation becomes increasingly sophisticated, understanding the rhetorical patterns in conspiratorial content is important for developing interventions such as targeted prebunking and assessing AI vulnerabilities. We introduce CONSPIRED (CONSPIR Evaluation Dataset), which captures the cognitive traits of conspiratorial ideation in multi-sentence excerpts (80-120 words) from online conspiracy articles, annotated using the CONSPIR cognitive framework. CONSPIRED is the first dataset of conspiratorial content annotated for general cognitive traits. Using CONSPIRED, we (i) develop computational models that identify conspiratorial traits and the dominant trait in text excerpts, and (ii) evaluate LLM robustness to conspiratorial inputs. We find that LLMs are readily misaligned by conspiratorial framing, reproducing its rhetorical patterns even when successfully deflecting comparable fact-checked misinformation.

阴谋论大模型安全认知分析数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。