arXiv:2502.19211cs.CL2025-02

发现部分大模型在否定信息后记忆减弱,提示其可能有类人记忆偏差。

Negation-Induced Forgetting in LLMs

  • 通过否定错误属性测试模型记忆,模拟人类认知现象
  • ChatGPT-3.5 明显出现否定导致遗忘,其他模型效果不一
  • 为理解大模型记忆机制提供首个实证线索,适合认知计算研究者

本研究探讨大型语言模型(LLMs)是否具备否定诱导遗忘(NIF)现象,即人类在否定某对象的错误属性后,对该对象的记忆会弱于肯定正确属性的情况(Mayo et al., 2014; Zang et al., 2023)。我们采用Zang等(2023)的实验框架,测试了ChatGPT-3.5、GPT-4o mini和LLaMA-3-70B-instruct。结果表明,ChatGPT-3.5表现出显著的NIF效应,否定信息比肯定信息更难被回忆;GPT-4o-mini呈现边缘显著的NIF;而LLaMA-3-70B未观察到该效应。研究首次为部分大模型中存在否定诱导遗忘提供了实证支持,暗示此类认知偏差可能在模型中自发产生,是理解模型记忆机制的初步探索。

原文摘要 · Abstract (English)

The study explores whether Large Language Models (LLMs) exhibit negation-induced forgetting (NIF), a cognitive phenomenon observed in humans where negating incorrect attributes of an object or event leads to diminished recall of this object or event compared to affirming correct attributes (Mayo et al., 2014; Zang et al., 2023). We adapted Zang et al. (2023) experimental framework to test this effect in ChatGPT-3.5, GPT-4o mini and Llama3-70b-instruct. Our results show that ChatGPT-3.5 exhibits NIF, with negated information being less likely to be recalled than affirmed information. GPT-4o-mini showed a marginally significant NIF effect, while LLaMA-3-70B did not exhibit NIF. The findings provide initial evidence of negation-induced forgetting in some LLMs, suggesting that similar cognitive biases may emerge in these models. This work is a preliminary step in understanding how memory-related phenomena manifest in LLMs.

大模型记忆认知偏差语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。