对比大模型与人类在歧义句理解上的表现差异。
When the LM misunderstood the human chuckled: Analyzing garden path effects in humans and language models
- 用心理语言学理论设计实验,测试人类与大模型对歧义句的理解。
- 发现部分大模型与人类在理解困难句上表现高度相似。
- 通过改写和图文生成任务验证了理解模式的一致性。
现代大语言模型(LLMs)在诸多语言任务中展现出类人能力,引发了对其与人类语言处理方式的比较研究兴趣。本文针对一种著名的语言难题——歧义句(garden-path constructions),开展人类与大模型的详细对比研究。基于心理语言学理论,我们提出关于歧义句难解原因的假设,并通过理解问题测试人类参与者及一系列大模型的表现。结果显示,人类与大模型均在特定句法复杂性面前表现困难,部分模型与人类的理解模式呈现显著相关性。为进一步验证,我们还对大模型在歧义句的改写与文本转图像任务中的表现进行测试,结果与句子理解任务一致,进一步支持了大模型在该类结构理解上的有限性与人类相似性。
原文摘要 · Abstract (English)
Modern Large Language Models (LLMs) have shown human-like abilities in many language tasks, sparking interest in comparing LLMs' and humans' language processing. In this paper, we conduct a detailed comparison of the two on a sentence comprehension task using garden-path constructions, which are notoriously challenging for humans. Based on psycholinguistic research, we formulate hypotheses on why garden-path sentences are hard, and test these hypotheses on human participants and a large suite of LLMs using comprehension questions. Our findings reveal that both LLMs and humans struggle with specific syntactic complexities, with some models showing high correlation with human comprehension. To complement our findings, we test LLM comprehension of garden-path constructions with paraphrasing and text-to-image generation tasks, and find that the results mirror the sentence comprehension question results, further validating our findings on LLM understanding of these constructions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。