用多轮角色化生成模拟假信息演化,揭示检测模型为何失效
MPCG: Multi-Round Persona-Conditioned Generation for Modeling the Evolution of Misinformation with LLMs
- 通过角色视角多轮迭代生成假信息,模拟其语言与价值观演变
- 生成内容认知负担更高,情感与道德框架完全贴合角色设定
- 暴露现有检测模型性能下降近50%,适合研究假信息演化者
假信息在传播中会动态变化,语言、表述方式和道德立场随受众调整。当前检测方法隐含假设假信息是静态的。本文提出MPCG——一种多轮、角色条件化的生成框架,模拟具有不同意识形态的代理对声明的逐轮重构。利用未受控的大语言模型(LLM)在多轮中生成角色专属声明,每一轮均基于前一轮输出进行条件生成,从而研究假信息的演化过程。我们通过人工标注、GPT-4o-mini标注、认知努力度量(可读性、困惑度)、情绪唤起度量(情感分析、道德性)、聚类分析、可行性评估及下游分类任务进行评估。结果显示,人类与GPT-4o-mini标注高度一致,流畅性判断差异较大。生成声明的认知负担高于原始声明,且持续体现角色对齐的情绪与道德框架。聚类与余弦相似度分析证实语义漂移的同时保持主题连贯性。可行性评估显示77%的生成声明具备实际应用潜力。分类结果表明,常见假信息检测器的宏平均F1值最高下降49.7%。代码已公开于https://github.com/bcjr1997/MPCG。
原文摘要 · Abstract (English)
Misinformation evolves as it spreads, shifting in language, framing, and moral emphasis to adapt to new audiences. However, current misinformation detection approaches implicitly assume that misinformation is static. We introduce MPCG, a multi-round, persona-conditioned framework that simulates how claims are iteratively reinterpreted by agents with distinct ideological perspectives. Our approach uses an uncensored large language model (LLM) to generate persona-specific claims across multiple rounds, conditioning each generation on outputs from the previous round, enabling the study of misinformation evolution. We evaluate the generated claims through human and LLM-based annotations, cognitive effort metrics (readability, perplexity), emotion evocation metrics (sentiment analysis, morality), clustering, feasibility, and downstream classification. Results show strong agreement between human and GPT-4o-mini annotations, with higher divergence in fluency judgments. Generated claims require greater cognitive effort than the original claims and consistently reflect persona-aligned emotional and moral framing. Clustering and cosine similarity analyses confirm semantic drift across rounds while preserving topical coherence. Feasibility results show a 77% feasibility rate, confirming suitability for downstream tasks. Classification results reveal that commonly used misinformation detectors experience macro-F1 performance drops of up to 49.7%. The code is available at https://github.com/bcjr1997/MPCG
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。