用对抗性反馈生成真实新闻的欺骗变体,评估大模型事实判断能力。
Real-time Factuality Assessment from Adversarial Feedback
- 通过RAG检测器反馈迭代改写新闻,生成挑战大模型的新样本。
- 使强RAG-GPT-4o检测器的ROC-AUC下降17.5个百分点。
- 适合关注大模型事实推理与安全评估的研究者。
我们发现,现有基于事实核查网站的新闻事实性评估方法,即使在大模型知识截止后仍能保持高准确率,这表明近期虚假信息因存在于预训练或检索语料中,或表现出显著但浅层的模式而易于识别。为此,我们提出一种新流程:利用基于RAG的检测器生成自然语言反馈,迭代改写实时新闻为具有欺骗性的变体,以测试模型对当前事件的推理能力。该方法使强RAG-GPT-4o检测器的二分类ROC-AUC绝对下降17.5个百分点。实验表明,RAG在生成与评估挑战性样本中起关键作用;无检索的模型易受未见事件和对抗攻击影响,而来自RAG评估的反馈能揭示更多欺骗模式。
原文摘要 · Abstract (English)
We show that existing evaluations for assessing the factuality of news from conventional sources, such as claims on fact-checking websites, result in high accuracies over time for LLM-based detectors-even after their knowledge cutoffs. This suggests that recent popular false information from such sources can be easily identified due to its likely presence in pre-training/retrieval corpora or the emergence of salient, yet shallow, patterns in these datasets. Instead, we argue that a proper factuality evaluation dataset should test a model's ability to reason about current events by retrieving and reading related evidence. To this end, we develop a novel pipeline that leverages natural language feedback from a RAG-based detector to iteratively modify real-time news into deceptive variants that challenge LLMs. Our iterative rewrite decreases the binary classification ROC-AUC by an absolute 17.5 percent for a strong RAG-based GPT-4o detector. Our experiments reveal the important role of RAG in both evaluating and generating challenging news examples, as retrieval-free LLM detectors are vulnerable to unseen events and adversarial attacks, while feedback from RAG-based evaluation helps discover more deceitful patterns.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。