测试大模型理解叙事时间意义的能力,发现其不如人类可靠。
How LLMs Comprehend Temporal Meaning in Narratives: A Case Study in Cognitive Evaluation of LLMs
- 用专家引导的探针方法评估模型对时态语义的理解
- 模型依赖典型性判断,因果推理能力弱,结果不一致
- 适合关注大模型认知局限与评测框架的研究者
大型语言模型(LLMs)展现出日益复杂的语言能力,但这些行为是反映类人认知,还是仅依赖模式识别仍存疑问。本研究考察了LLMs在以往用于人类实验的叙事时态语义处理能力。通过专家参与的探针管道,我们开展系列针对性实验,评估模型是否以类人方式构建语义表征和语用推断。结果表明,LLMs过度依赖典型性,产生不一致的时态判断,且难以从时态中进行因果推理,提示其对叙事的理解本质不同于人类,缺乏稳健的叙事理解能力。此外,我们构建了一个标准化实验框架,可用于可靠评估LLMs的认知与语言能力。
原文摘要 · Abstract (English)
Large language models (LLMs) exhibit increasingly sophisticated linguistic capabilities, yet the extent to which these behaviors reflect human-like cognition versus advanced pattern recognition remains an open question. In this study, we investigate how LLMs process the temporal meaning of linguistic aspect in narratives that were previously used in human studies. Using an Expert-in-the-Loop probing pipeline, we conduct a series of targeted experiments to assess whether LLMs construct semantic representations and pragmatic inferences in a human-like manner. Our findings show that LLMs over-rely on prototypicality, produce inconsistent aspectual judgments, and struggle with causal reasoning derived from aspect, raising concerns about their ability to fully comprehend narratives. These results suggest that LLMs process aspect fundamentally differently from humans and lack robust narrative understanding. Beyond these empirical findings, we develop a standardized experimental framework for the reliable assessment of LLMs' cognitive and linguistic capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。