构建首个大规模谣言评估基准,揭示大模型在谣言识别中的脆弱性。
How does Misinformation Affect Large Language Model Behaviors and Preferences?
- 构建含超1030万条谣言的综合评测集,涵盖知识冲突与风格差异
- 实测显示大模型虽能辨识谣言,仍易受内容冲突和表达风格干扰
- 提出RtD新方法提升检测能力,适合安全、可信AI研究者使用
大型语言模型(LLMs)在知识密集型任务中表现卓越,但在面对虚假信息时仍显脆弱。现有研究多关注模型对抗虚假信息的能力,却缺乏对模型行为与知识偏好受虚假信息影响程度的细粒度分析。为此,本文提出MisBench——当前最大且最全面的评估基准,用于衡量LLMs在虚假信息上的行为与知识偏好。MisBench包含10,346,712条虚假信息,首次同时考虑知识层面的矛盾与风格上的变异。实证结果表明,尽管大模型具备一定的虚假信息辨别能力,但仍易受知识冲突与风格变化的影响。基于此发现,我们进一步提出一种名为Reconstruct to Discriminate(RtD)的新方法,以增强模型对虚假信息的检测能力。本研究为理解大模型与虚假信息的交互提供了重要洞见,我们相信MisBench可作为评估基于大模型的检测器的有效基准,提升其在真实场景中的可靠性。代码与数据已开源:https://github.com/GKNL/MisBench。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have shown remarkable capabilities in knowledge-intensive tasks, while they remain vulnerable when encountering misinformation. Existing studies have explored the role of LLMs in combating misinformation, but there is still a lack of fine-grained analysis on the specific aspects and extent to which LLMs are influenced by misinformation. To bridge this gap, we present MisBench, the current largest and most comprehensive benchmark for evaluating LLMs' behavior and knowledge preference toward misinformation. MisBench consists of 10,346,712 pieces of misinformation, which uniquely considers both knowledge-based conflicts and stylistic variations in misinformation. Empirical results reveal that while LLMs demonstrate comparable abilities in discerning misinformation, they still remain susceptible to knowledge conflicts and stylistic variations. Based on these findings, we further propose a novel approach called Reconstruct to Discriminate (RtD) to strengthen LLMs' ability to detect misinformation. Our study provides valuable insights into LLMs' interactions with misinformation, and we believe MisBench can serve as an effective benchmark for evaluating LLM-based detectors and enhancing their reliability in real-world applications. Codes and data are available at https://github.com/GKNL/MisBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。