研究大模型如何处理多语言语法歧义,发现对亚洲语言处理较差
Multilingual Relative Clause Attachment Ambiguity Resolution in Large Language Models
- 测试大模型在多语言中解析相对从句歧义的能力
- 欧语系表现良好,日韩语常误译为英文
- 揭示模型对非欧洲语言的处理短板,适合语言模型优化者参考
本研究考察大型语言模型(LLMs)如何解决相对从句(RC)附着歧义,并对比其与人类句法处理的表现。聚焦相对从句长度和复杂限定词短语(DP)的句法位置两个语言因素,评估大模型在语言复杂性下的理解能力。实验涵盖英语、西班牙语、法语、德语、日语和韩语,使用Claude、Gemini和Llama等多个模型。结果显示,这些模型在印欧语系(英语、西班牙语、法语、德语)中表现良好,但在日语和韩语中常出现错误,倾向于将结果错误翻译成英文。研究揭示了大模型在处理语言歧义时的不一致性,尤其在非欧洲语言中表现不足,强调需改进模型设计以提升跨语言准确性和类人处理能力。
原文摘要 · Abstract (English)
This study examines how large language models (LLMs) resolve relative clause (RC) attachment ambiguities and compares their performance to human sentence processing. Focusing on two linguistic factors, namely the length of RCs and the syntactic position of complex determiner phrases (DPs), we assess whether LLMs can achieve human-like interpretations amid the complexities of language. In this study, we evaluated several LLMs, including Claude, Gemini and Llama, in multiple languages: English, Spanish, French, German, Japanese, and Korean. While these models performed well in Indo-European languages (English, Spanish, French, and German), they encountered difficulties in Asian languages (Japanese and Korean), often defaulting to incorrect English translations. The findings underscore the variability in LLMs' handling of linguistic ambiguities and highlight the need for model improvements, particularly for non-European languages. This research informs future enhancements in LLM design to improve accuracy and human-like processing in diverse linguistic environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。