arXiv:2409.15334cs.CL2024-09被引 3

测试大模型对西班牙语的理解能力,发现语法仍不如母语者。

Evaluating Large Language Models with Tests of Spanish as a Foreign Language: Pass or Fail?

  • 用西班牙语考试题集TELEIA评估大模型语言理解能力
  • 模型阅读理解和词义分析表现良好,但语法错误率高
  • 适合关注多语言大模型能力边界的研究者

大型语言模型(LLMs)在英文问答和自然语言理解任务上已有大量评估,但多数用户并非英语母语者。因此,研究其在不同语言层级(从段落到词素)的理解能力至关重要。本文使用TELEIA这一近期发布的基准测试,该测试题型与面向非母语者的西班牙语考试相似,涵盖阅读理解、构词法、词义与组合语义、语法等内容。结果表明,尽管大模型在理解西班牙语方面表现良好,但在语法掌握程度上仍远未达到母语者的水平。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have been profusely evaluated on their ability to answer questions on many topics and their performance on different natural language understanding tasks. Those tests are usually conducted in English, but most LLM users are not native English speakers. Therefore, it is of interest to analyze how LLMs understand other languages at different levels: from paragraphs to morphems. In this paper, we evaluate the performance of state-of-the-art LLMs in TELEIA, a recently released benchmark with similar questions to those of Spanish exams for foreign students, covering topics such as reading comprehension, word formation, meaning and compositional semantics, and grammar. The results show that LLMs perform well at understanding Spanish but are still far from achieving the level of a native speaker in terms of grammatical competence.

大模型评估多语言西班牙语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。