研究大模型如何判断带不确定性的假信息,发现25%的假话被误判为非假。
On Fact and Frequency: LLM Responses to Misinformation Expressed with Uncertainty
- 将已知假话改写为不确定表述,测试大模型判断变化
- 25%假话在不确定性表达后被判定为非假
- 模型对频率估计与真假判断存在微弱相关
我们研究大语言模型对带有不确定性的虚假信息的判断。实验考察了三种主流大模型(GPT-4o、LlaMA3、DeepSeek-v2)在将已验证为假的陈述按不确定性类型改写后的响应。结果显示,经过改写后,大模型将这些陈述从‘假’改为‘非假’的比例达25%。分析表明,这一变化无法用人类敏感的模态、语言线索或论证策略解释,唯一例外是信念类改写(如“据信……”这类表达)。为进一步理解,我们让模型对这些不确定陈述在人群中出现的频率进行估计。结果发现,模型对频率的估计与真假判断之间存在微弱但显著的相关性。
原文摘要 · Abstract (English)
We study LLM judgments of misinformation expressed with uncertainty. Our experiments study the response of three widely used LLMs (GPT-4o, LlaMA3, DeepSeek-v2) to misinformation propositions that have been verified false and then are transformed into uncertain statements according to an uncertainty typology. Our results show that after transformation, LLMs change their factchecking classification from false to not-false in 25% of the cases. Analysis reveals that the change cannot be explained by predictors to which humans are expected to be sensitive, i.e., modality, linguistic cues, or argumentation strategy. The exception is doxastic transformations, which use linguistic cue phrases such as "It is believed ...".To gain further insight, we prompt the LLM to make another judgment about the transformed misinformation statements that is not related to truth value. Specifically, we study LLM estimates of the frequency with which people make the uncertain statement. We find a small but significant correlation between judgment of fact and estimation of frequency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。