arXiv:2409.15454cs.CLcs.AI2024-09EMNLP被引 9

大模型在简单任务中表现好,但面对细微变化时会犯婴儿级的固执错误。

In-Context Learning May Not Elicit Trustworthy Reasoning: A-Not-B Errors in Pretrained Language Models

论文配图:In-Context Learning May Not Elicit Trustworthy Reasoning: A-Not-B Errors in Pretrained Language Models
图 1 · 摘自论文原文
  • 设计类比婴儿实验的文本问答任务,测试模型抑制旧习惯的能力。
  • 模型在上下文微调后错误率最高上升83.3%,推理能力大幅下降。
  • 揭示当前大模型缺乏真正的认知灵活性,适合研究认知局限的学者参考。

近年来人工智能发展催生出能以类人方式执行任务的大型语言模型(LLMs),但在某些领域仅具备婴儿级认知能力。其中一种表现是A-Not-B错误——即使观察到条件已变,仍重复先前受奖励的行为,体现其缺乏抑制控制能力。本文设计了一种基于文本的多选题问答场景,模拟经典A-Not-B实验,系统测试LLMs的抑制控制能力。结果发现,先进模型(如Llama3-8b)在使用上下文学习(ICL)时表现稳定,但当上下文发生细微变化时,推理任务错误率显著上升,最高达83.3%。这表明,此类模型在此类任务中的抑制控制能力仅相当于人类婴儿,常无法抑制先前建立的反应模式。

原文摘要 · Abstract (English)

Recent advancements in artificial intelligence have led to the creation of highly capable large language models (LLMs) that can perform tasks in a human-like manner. However, LLMs exhibit only infant-level cognitive abilities in certain areas. One such area is the A-Not-B error, a phenomenon seen in infants where they repeat a previously rewarded behavior despite well-observed changed conditions. This highlights their lack of inhibitory control -- the ability to stop a habitual or impulsive response. In our work, we design a text-based multi-choice QA scenario similar to the A-Not-B experimental settings to systematically test the inhibitory control abilities of LLMs. We found that state-of-the-art LLMs (like Llama3-8b) perform consistently well with in-context learning (ICL) but make errors and show a significant drop of as many as 83.3% in reasoning tasks when the context changes trivially. This suggests that LLMs only have inhibitory control abilities on par with human infants in this regard, often failing to suppress the previously established response pattern during ICL.

认知局限抑制控制大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。