对比人类与大模型在概念和指称信息干扰下的处理差异
Distinct dynamics of conceptual and referential disruptions in human reading and large language model processing

- 通过干扰叙事中的概念或指称信息,追踪阅读与模型处理过程
- 概念干扰导致局部强反应,指称干扰影响更分散且持续更久
- 大模型的输出层显示指称干扰初始影响更大,符合人类阅读规律
语言意义基于概念内容,词语进入语篇后产生对特定实体的指称。为考察这两类意义的处理动态,我们对短篇叙事中的概念或指称信息进行选择性干扰,追踪其在人类自定速阅读和大语言模型预测与表征过程中的影响。人类阅读中,概念干扰引发强烈但局部的加工成本,出现在扰动词后立即出现并迅速达到峰值,随后快速下降;指称干扰效应较弱,随后续词逐渐减弱,且受句界调节更强。在语言模型中,两类干扰均在操作词处立即显现。上下文模型意外度(surprisal)模式与人类阅读高度一致:概念干扰产生更大、更局部集中且快速衰减的效应,而指称干扰则产生较小但更渐进的下游影响。输出层表示则呈现不同模式:指称干扰引发更大的初始偏移,而两种干扰后续均表现为幂律衰减。这些结果共同支持两类意义处理动态可区分:概念信息带来更局部集中的整合代价,而指称信息则涉及更分布式的语篇层面身份维持过程。
原文摘要 · Abstract (English)
Linguistic meaning is grounded in conceptual content, from which reference to particular entities emerges as words enter discourse. To examine the processing dynamics associated with these two dimensions of meaning, we selectively disrupted conceptual or referential information in short narratives and traced the resulting effects in human self-paced reading and in the predictive and representational processing of large language models. In human reading, conceptual disruptions produced a strong but localized processing cost, emerging immediately after the distorted word, reaching an early maximum, and then declining rapidly. Referential disruptions produced weaker effects, which decreased more gradually across subsequent words, and were more strongly modulated by sentence boundaries. In the language model, both disruptions emerged immediately at the manipulated word. Contextual model surprisal showed a pattern closely paralleling human reading: conceptual disruption produced a larger, more locally concentrated effect that decayed rapidly, whereas referential disruption produced a smaller and more gradual downstream effect. Output-layer representations showed a different pattern: referential disruption produced a larger initial displacement, while both distortions were subsequently characterized by power-law decay. Together, these results provide convergent evidence for distinguishable processing dynamics of two types of meaning: conceptual information imposes a more locally concentrated integration cost, whereas referential information engages a more distributed process of maintaining discourse-level identity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。