首个隐性错误信息基准测试发现大模型极易受误导,需警惕潜在风险。
How to Protect Yourself from 5G Radiation? Investigating LLM Responses to Implicit Misinformation
- 构建首个隐性错误信息基准EchoMist,模拟真实对话中的隐蔽错误前提。
- 15个主流大模型在该任务中表现差,多数无法识别错误前提并生成反事实解释。
- 自警与RAG方法有一定缓解作用,但隐性误导仍是重大安全挑战。
随着大语言模型(LLMs)在各类场景中广泛应用,其可能隐性传播错误信息的问题成为关键安全关切。现有研究主要评估显式虚假陈述,忽略了错误信息常以未被质疑的前提形式悄然出现。本文构建了首个全面的隐性错误信息基准EchoMist,其中错误假设嵌入至对模型的提问中,涵盖来自现实人机对话与社交媒体的广泛、有害且持续演化的隐性错误信息。对15个先进大模型的实证研究表明,当前模型在此任务上表现令人担忧,往往无法识别错误前提,并生成反事实解释。我们还探究了两种缓解策略——自警机制与检索增强生成(RAG),结果表明尽管有所改善,但隐性错误信息仍是持久挑战,凸显防范此类风险的紧迫性。
原文摘要 · Abstract (English)
As Large Language Models (LLMs) are widely deployed in diverse scenarios, the extent to which they could tacitly spread misinformation emerges as a critical safety concern. Current research primarily evaluates LLMs on explicit false statements, overlooking how misinformation often manifests subtly as unchallenged premises in real-world interactions. We curated EchoMist, the first comprehensive benchmark for implicit misinformation, where false assumptions are embedded in the query to LLMs. EchoMist targets circulated, harmful, and ever-evolving implicit misinformation from diverse sources, including realistic human-AI conversations and social media interactions. Through extensive empirical studies on 15 state-of-the-art LLMs, we find that current models perform alarmingly poorly on this task, often failing to detect false premises and generating counterfactual explanations. We also investigate two mitigation methods, i.e., Self-Alert and RAG, to enhance LLMs' capability to counter implicit misinformation. Our findings indicate that EchoMist remains a persistent challenge and underscore the critical need to safeguard against the risk of implicit misinformation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。