arXiv:2410.13392cs.CL2024-10被引 6

大模型无法像人一样准确预判自己能否记住信息,暴露了其元认知短板。

Judgment of Learning: A Human Ability Beyond Generative Artificial Intelligence

  • 用跨主体预测模型对比人类与大模型对记忆的自我判断能力。
  • 人类能准确预测记忆表现,但GPT-3.5-turbo、GPT-4-turbo、GPT-4o均无此能力。
  • 在语义契合或不符情境下,模型始终无法匹配人类的元认知表现。

大型语言模型(LLMs)在多种语言任务中日益逼近人类认知,但其在元认知——尤其是预测记忆表现——方面的能力仍未知。本文引入跨主体预测模型,评估基于ChatGPT的LLMs是否与人类的判断学习(JOL)一致,即个体对自己未来记忆表现的自我预测。我们测试了人类与LLMs对句子对的记忆判断,其中包含花园路径句(需重新分析的误导性句子),通过操控上下文契合度(契合/不契合)探究内在线索(如语义相关性)如何影响人类与大模型的JOL。结果表明,人类的JOL可有效预测实际记忆表现,而所有测试的LLMs(GPT-3.5-turbo、GPT-4-turbo、GPT-4o)均未表现出类似预测准确性。该差异在契合与不契合语境下均存在。研究揭示:尽管大模型能在对象层面模拟人类认知,但在元认知层面仍显不足,难以捕捉个体记忆预测的变异性。这一发现凸显了提升大模型自监控能力的必要性,有助于教育应用、个性化学习及人机交互中降低对人工监督的依赖,推动更自主、无缝的智能系统集成。

原文摘要 · Abstract (English)

Large language models (LLMs) increasingly mimic human cognition in various language-based tasks. However, their capacity for metacognition - particularly in predicting memory performance - remains unexplored. Here, we introduce a cross-agent prediction model to assess whether ChatGPT-based LLMs align with human judgments of learning (JOL), a metacognitive measure where individuals predict their own future memory performance. We tested humans and LLMs on pairs of sentences, one of which was a garden-path sentence - a sentence that initially misleads the reader toward an incorrect interpretation before requiring reanalysis. By manipulating contextual fit (fitting vs. unfitting sentences), we probed how intrinsic cues (i.e., relatedness) affect both LLM and human JOL. Our results revealed that while human JOL reliably predicted actual memory performance, none of the tested LLMs (GPT-3.5-turbo, GPT-4-turbo, and GPT-4o) demonstrated comparable predictive accuracy. This discrepancy emerged regardless of whether sentences appeared in fitting or unfitting contexts. These findings indicate that, despite LLMs' demonstrated capacity to model human cognition at the object-level, they struggle at the meta-level, failing to capture the variability in individual memory predictions. By identifying this shortcoming, our study underscores the need for further refinements in LLMs' self-monitoring abilities, which could enhance their utility in educational settings, personalized learning, and human-AI interactions. Strengthening LLMs' metacognitive performance may reduce the reliance on human oversight, paving the way for more autonomous and seamless integration of AI into tasks requiring deeper cognitive awareness.

元认知大模型记忆预测人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。