对比人类与大模型在语法歧义解析中对世界知识的依赖差异
Who Relies More on World Knowledge and Bias for Syntactic Ambiguity Resolution: Humans or LLMs?
- 构建多语言歧义句数据集MultiWho,评估相对子句附着偏好
- 大模型始终倾向局部附着,对语言差异响应弱,准确率仅67%
- 适合研究大模型语言理解局限性及多语言训练改进
本研究探讨近期大语言模型(LLMs)在六种类型多样的语言(英语、汉语、日语、韩语、俄语、西班牙语)中如何处理相对子句附着歧义,并利用世界知识偏见进行消歧。我们创建了新型数据集MultiWho,用于细粒度评估歧义与非歧义情境下的相对子句附着偏好。三款大模型实验表明,与人类相反,大模型始终表现出对局部附着的偏好,对句法变化或语言特异性附着模式响应有限。尽管在非歧义情况下表现良好,但大模型僵化地优先使用世界知识偏见,缺乏人类语言处理的灵活性。研究强调需采用更多样化、具有语用细微差别的多语言训练,以提升大模型对复杂结构的处理能力与类人理解水平。
原文摘要 · Abstract (English)
This study explores how recent large language models (LLMs) navigate relative clause attachment {ambiguity} and use world knowledge biases for disambiguation in six typologically diverse languages: English, Chinese, Japanese, Korean, Russian, and Spanish. We describe the process of creating a novel dataset -- MultiWho -- for fine-grained evaluation of relative clause attachment preferences in ambiguous and unambiguous contexts. Our experiments with three LLMs indicate that, contrary to humans, LLMs consistently exhibit a preference for local attachment, displaying limited responsiveness to syntactic variations or language-specific attachment patterns. Although LLMs performed well in unambiguous cases, they rigidly prioritized world knowledge biases, lacking the flexibility of human language processing. These findings highlight the need for more diverse, pragmatically nuanced multilingual training to improve LLMs' handling of complex structures and human-like comprehension.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。