arXiv:2505.22354cs.CL2025-05被引 12

大模型在高风险假前提下难识别误导信息

LLMs Struggle to Reject False Presuppositions when Misinformation Stakes are High

  • 通过语言预设分析测试模型对虚假前提的敏感度
  • 三款模型在政治语境中识别率均低于60%
  • 适合关注生成式AI misinformation风险的研究者

本文研究大模型在面对虚假预设时的表现,以及语言特征如何影响其对错误预设内容的回应。预设以隐含方式将信息视为既定事实,极易嵌入争议或虚假信息,引发对大模型是否如人类一样,在高风险虚假信息情境下仍无法识别并纠正误导性假设的担忧。基于语言预设分析方法,我们系统考察了模型在不同条件下的敏感性,重点关注政治语境中语言结构、政党立场及情景概率等因素的影响。使用新构建的数据集,评估了GPT-4-o、LLama-3-8B和Mistral-7B-v03三款模型。结果表明,模型普遍难以识别虚假预设,表现因条件而异。该研究证明,语言预设分析是揭示大模型响应中政治谣言强化机制的有效工具。

原文摘要 · Abstract (English)

This paper examines how LLMs handle false presuppositions and whether certain linguistic factors influence their responses to falsely presupposed content. Presuppositions subtly introduce information as given, making them highly effective at embedding disputable or false information. This raises concerns about whether LLMs, like humans, may fail to detect and correct misleading assumptions introduced as false presuppositions, even when the stakes of misinformation are high. Using a systematic approach based on linguistic presupposition analysis, we investigate the conditions under which LLMs are more or less sensitive to adopt or reject false presuppositions. Focusing on political contexts, we examine how factors like linguistic construction, political party, and scenario probability impact the recognition of false presuppositions. We conduct experiments with a newly created dataset and examine three LLMs: OpenAI's GPT-4-o, Meta's LLama-3-8B, and MistralAI's Mistral-7B-v03. Our results show that the models struggle to recognize false presuppositions, with performance varying by condition. This study highlights that linguistic presupposition analysis is a valuable tool for uncovering the reinforcement of political misinformation in LLM responses.

大模型安全虚假信息预设分析政治偏见

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。