arXiv:2505.14917cs.CL2025-05被引 2

针对大模型伪装阴谋论的新检测方法,提升识别能力。

ConspEmoLLM-v2: A robust and stable model to detect sentiment-transformed conspiracy theories

  • 用大模型重写阴谋论文本降低负面情绪,构建新数据集
  • 新模型在原始文本上性能持平,伪装文本上显著更优
  • 适合研究虚假信息检测与大模型安全的开发者

尽管大语言模型(LLMs)带来诸多好处,也可能造成危害,如自动生成虚假信息,包括阴谋论。此外,这些模型可通过改变典型文本特征(如将强烈负面情绪转为更积极语气)来伪装阴谋论。现有检测方法多基于人工撰写的文本训练,其特征与大模型生成内容存在差异。先前的ConspEmoLLM模型依赖人类撰写阴谋论的典型情绪特征,易被刻意伪装的内容绕过。为此,我们首先构建了增强版的ConDID数据集——ConDID-v2,将原有人工撰写阴谋论推文用大模型改写为情绪更温和的版本,并通过人工与大模型联合评估验证改写质量。随后使用ConDID-v2训练了ConspEmoLLM-v2,作为ConspEmoLLM的升级版。实验表明,ConspEmoLLM-v2在原有人工文本上性能持平或超越,而在处理情感伪装后的推文时,显著优于ConspEmoLLM及多个基线模型。

原文摘要 · Abstract (English)

Despite the many benefits of large language models (LLMs), they can also cause harm, e.g., through automatic generation of misinformation, including conspiracy theories. Moreover, LLMs can also ''disguise'' conspiracy theories by altering characteristic textual features, e.g., by transforming their typically strong negative emotions into a more positive tone. Although several studies have proposed automated conspiracy theory detection methods, they are usually trained using human-authored text, whose features can vary from LLM-generated text. Furthermore, several conspiracy detection models, including the previously proposed ConspEmoLLM, rely heavily on the typical emotional features of human-authored conspiracy content. As such, intentionally disguised content may evade detection. To combat such issues, we firstly developed an augmented version of the ConDID conspiracy detection dataset, ConDID-v2, which supplements human-authored conspiracy tweets with versions rewritten by an LLM to reduce the negativity of their original sentiment. The quality of the rewritten tweets was verified by combining human and LLM-based assessment. We subsequently used ConDID-v2 to train ConspEmoLLM-v2, an enhanced version of ConspEmoLLM. Experimental results demonstrate that ConspEmoLLM-v2 retains or exceeds the performance of ConspEmoLLM on the original human-authored content in ConDID, and considerably outperforms both ConspEmoLLM and several other baselines when applied to sentiment-transformed tweets in ConDID-v2. The project will be available at https://github.com/lzw108/ConspEmoLLM.

阴谋论检测大模型安全情感伪装文本对抗

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。