arXiv:2501.18649cs.CLcs.AI2025-01被引 13

检测大模型生成的假新闻,发现改写后更难识别

Fake News Detection After LLM Laundering: Measurement and Explanation

  • 用大模型改写假新闻,测试检测器效果
  • 改写后假新闻检测准确率下降超过20%
  • 发现情感变化是误判主因,适合安全与检测研究者

大型语言模型(LLMs)具备生成高度可信且语境相关的假新闻的能力,可能加剧虚假信息传播。尽管已有大量关于人工撰写文本的假新闻检测研究,但针对大模型生成假新闻的检测仍处于探索阶段。本研究评估检测器识别大模型改写假新闻的有效性,特别考察在检测流程中加入改写步骤是否有助于或阻碍检测。研究发现:(1) 检测器对大模型改写的假新闻识别能力显著弱于对人工撰写文本的识别;(2) 不同模型在逃避检测、改写以规避检测及保持语义相似性方面表现各异;(3) 通过LIME解释分析,发现检测失败可能源于情感偏移;(4) 发现一个令人担忧的趋势:即使BERTScore很高,样本仍可能出现情感偏移;(5) 提供一对数据集,扩充现有数据集并包含改写输出与评分。数据集已公开于GitHub。

原文摘要 · Abstract (English)

With their advanced capabilities, Large Language Models (LLMs) can generate highly convincing and contextually relevant fake news, which can contribute to disseminating misinformation. Though there is much research on fake news detection for human-written text, the field of detecting LLM-generated fake news is still under-explored. This research measures the efficacy of detectors in identifying LLM-paraphrased fake news, in particular, determining whether adding a paraphrase step in the detection pipeline helps or impedes detection. This study contributes: (1) Detectors struggle to detect LLM-paraphrased fake news more than human-written text, (2) We find which models excel at which tasks (evading detection, paraphrasing to evade detection, and paraphrasing for semantic similarity). (3) Via LIME explanations, we discovered a possible reason for detection failures: sentiment shift. (4) We discover a worrisome trend for paraphrase quality measurement: samples that exhibit sentiment shift despite a high BERTSCORE. (5) We provide a pair of datasets augmenting existing datasets with paraphrase outputs and scores. The dataset is available on GitHub

假新闻检测大模型安全情感偏移可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。