测试5个大模型对指代消解的人类认知规律拟合度
Human-Like Anaphor Resolution in Large Language Models
- 用人类阅读时间建模,对比模型困惑度与人类表现
- 部分模型对上下文显著性和距离敏感,但对语义干扰不敏感
- 揭示大模型在哪些情况下像人、哪些不像,适合做语言理解研究
指代词是引用其他表达(即先行词)的词语,连接两者的过程称为指代消解。认知科学发现,指代消解的速度和成功率受话语结构、情境模型特性及语义因素影响。本文研究五个具有开放权重的大语言模型(GPT-2-XL、Llama-3.1-8B、Pythia-12B、Mistral-7B 和 Mistral-24B)是否也受到这些因素影响。为建模处理难度,采用标准链接假说,将人类阅读时间与模型在指代词处的困惑度(surprisal)关联;作为第二项行为指标,比较模型在探测指代词先行词的理解题上的准确率与人类表现。结果表明存在选择性认知对齐:部分模型在指代消解中表现出对话语显著性和距离因素的人类级敏感性,但在语义干扰效应上敏感性较弱或缺失。该发现限定了大模型模拟人类指代消解的适用条件。
原文摘要 · Abstract (English)
Anaphors are expressions that refer to other expressions, called antecedents. The process of connecting the two is called resolution. Cognitive science has identified multiple factors that affect the speed and success of anaphor resolution, including discourse structure, situation-model properties, and semantic factors. Here, we investigate whether these factors also affect anaphor resolution in five Large Language Models (LLMs) with open weights: GPT-2-XL, Llama-3.1-8B, Pythia-12B, Mistral-7B, and Mistral-24B. To model processing difficulty, we adopt the standard linking hypothesis that relates human reading times to model surprisal at the anaphor. As a second behavioral measure, we compare model accuracy to human accuracy on comprehension questions probing the antecedents of anaphors. The results show selective cognitive alignment: some LLMs exhibit human-like sensitivity to discourse prominence and distance-based factors in anaphor resolution, while showing weaker or absent sensitivity to semantic interference effects. These findings delimit the conditions under which LLMs approximate human anaphor resolution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。