对比人类与大模型在复杂句式上的理解难度,发现两者在迂回句上表现差异显著。
Comparing Human and Language Models Sentence Processing Difficulties on Complex Structures
- 统一框架下测试人类与5类大模型对7种复杂句式的理解能力。
- 最强模型(GPT-5)在非迂回句准确率达93.7%,但在迂回句仅46.8%。
- 模型参数量越大,其句式理解排名与人类越接近,但过弱或过强模型均失真。
大型语言模型(LLMs)已能流利与人类对话,但它们是否经历类似人类的语义处理困难?我们系统比较了人类与五类前沿大模型在七种挑战性语言结构上的句子理解能力。在统一实验框架下,收集了人类及各模型的句子理解数据,涵盖不同规模与训练方式。结果显示,尽管大模型整体在目标结构上表现不佳,尤其在迂回句(GP)上——最强模型GPT-5在非迂回句准确率达93.7%,但在迂回句仅46.8%。同时,随着参数量增加,模型与人类在句式理解排名上的相关性上升。针对每种结构,还测试了无复杂结构的对照句,发现人类中存在性能差距的现象同样出现在大模型中,但有两个例外:过弱模型在两类句子上表现均低,过强模型则表现均高。这些结果揭示了人类与大模型在句理解上的趋同与分化,为理解两者相似性提供了新视角。
原文摘要 · Abstract (English)
Large language models (LLMs) that fluently converse with humans are a reality - but do LLMs experience human-like processing difficulties? We systematically compare human and LLM sentence comprehension across seven challenging linguistic structures. We collect sentence comprehension data from humans and five families of state-of-the-art LLMs, varying in size and training procedure in a unified experimental framework. Our results show LLMs overall struggle on the target structures, but especially on garden path (GP) sentences. Indeed, while the strongest models achieve near perfect accuracy on non-GP structures (93.7% for GPT-5), they struggle on GP structures (46.8% for GPT-5). Additionally, when ranking structures based on average performance, rank correlation between humans and models increases with parameter count. For each target structure, we also collect data for their matched baseline without the difficult structure. Comparing performance on the target vs. baseline sentences, the performance gap observed in humans holds for LLMs, with two exceptions: for models that are too weak performance is uniformly low across both sentence types, and for models that are too strong the performance is uniformly high. Together, these reveal convergence and divergence in human and LLM sentence comprehension, offering new insights into the similarity of humans and LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。