测试大模型能否理解场景中元素与上下文的关系,发现主要靠记忆而非理解。
Memorization $\neq$ Understanding: Do Large Language Models Have the Ability of Scenario Cognition?
- 构建场景数据集,从输出和内部表征双角度评估模型理解能力。
- 模型在简单场景中仍无法稳定关联元素与上下文,依赖表面记忆。
- 适合研究大模型语义理解局限性和认知机制的学者参考。
基于海量多样的文本数据,大型语言模型(LLMs)在众多自然语言处理任务中表现出色。然而,一个核心问题依然存在:其泛化能力是源于对训练数据的机械记忆,还是深层语义理解?为此,我们提出一种双视角评估框架,用于检验 LLMs 的场景认知能力——即在上下文中将语义场景元素与其论据关联的能力。具体地,我们构建了一个包含多样虚构事实描述的场景数据集,并标注了场景元素。通过评估模型回答场景相关问题的能力(输出视角)以及探测其内部表示中是否编码了场景元素-论据关联(内部表示视角),实验发现当前 LLMs 主要依赖浅层记忆,在简单案例中也难以实现稳健的语义场景认知。这些结果揭示了 LLMs 在语义理解上的关键局限,并为提升其认知能力提供了认知洞察。
原文摘要 · Abstract (English)
Driven by vast and diverse textual data, large language models (LLMs) have demonstrated impressive performance across numerous natural language processing (NLP) tasks. Yet, a critical question persists: does their generalization arise from mere memorization of training data or from deep semantic understanding? To investigate this, we propose a bi-perspective evaluation framework to assess LLMs' scenario cognition - the ability to link semantic scenario elements with their arguments in context. Specifically, we introduce a novel scenario-based dataset comprising diverse textual descriptions of fictional facts, annotated with scenario elements. LLMs are evaluated through their capacity to answer scenario-related questions (model output perspective) and via probing their internal representations for encoded scenario elements-argument associations (internal representation perspective). Our experiments reveal that current LLMs predominantly rely on superficial memorization, failing to achieve robust semantic scenario cognition, even in simple cases. These findings expose critical limitations in LLMs' semantic understanding and offer cognitive insights for advancing their capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。