arXiv:2601.03981cs.CL2026-01

用对抗反馈提升假新闻检测模型鲁棒性,效果优于主流方法。

RADAR: Retrieval-Augmented Detector with Adversarial Refinement for Robust Fake News Detection

  • 通过改写真实文章生成逼真假新闻,配合检索增强检测器。
  • 在基准测试中超越强基线模型,检索能力提升贡献最大。
  • 适合需要高鲁棒性的假新闻检测场景,尤其对抗未知攻击者。

为有效应对大模型生成的虚假信息传播,我们提出RADAR——一种带有对抗精炼的检索增强检测框架。该方法使用生成器对真实文章进行事实扰动改写,搭配轻量级检测器,利用密集段落检索验证内容真实性。为实现有效协同进化,引入自然语言形式的对抗反馈(VAF),以结构化文本批评指导生成器生成更复杂欺骗性内容,推动检测器持续优化。在假新闻检测基准上,RADAR持续优于多个强基线与通用大模型结合检索的方法。分析显示,检测端的检索带来最大性能提升,而VAF与少量示例演示提供互补增益。此外,RADAR在未见过的外部攻击者生成的假新闻上仍具更好迁移能力,表明其具备超越训练环境的鲁棒性。

原文摘要 · Abstract (English)

To efficiently combat the spread of LLM-generated misinformation, we present RADAR, a Retrieval-Augmented Detector with Adversarial Refinement for robust fake news detection. Our approach employs a generator that rewrites real articles with factual perturbations, paired with a lightweight detector that verifies claims using dense passage retrieval. To enable effective co-evolution, we introduce verbal adversarial feedback (VAF). Rather than relying on scalar rewards, VAF issues structured natural-language critiques; these guide the generator toward more sophisticated evasion attempts, compelling the detector to adapt and improve. On a fake news detection benchmark, RADAR consistently outperforms strong retrieval-augmented trainable baselines, as well as general-purpose LLMs with retrieval. Further analysis shows that detector-side retrieval yields the largest gains, while VAF and few-shot demonstrations provide complementary benefits. RADAR also transfers better to fake news generated by an unseen external attacker, indicating improved robustness beyond the co-evolved training setting.

假新闻检测对抗训练检索增强大模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。