arXiv:2503.00995cs.CL2025-03被引 6

测试大模型对波兰语言文化理解能力,发现普遍短板

Evaluating Polish linguistic and cultural competency in large language models

  • 构建600题波兰文化专项评测集,涵盖历史、地理、艺术等六类
  • 超30个开源与商用模型参与测试,揭示对波兰文化认知不足
  • 适合关注多语言文化理解的AI研究者与开发者参考

大型语言模型在多语言文本处理与生成方面日益成熟,有助于解决现实问题。然而,语言理解远不止文本分析,还需掌握文化背景,包括日常生活、历史事件、传统习俗、民间传说、文学及流行文化等。缺乏此类知识可能导致误解和难以察觉的错误。为评估模型对波兰文化背景的理解能力,我们构建了包含600道人工设计题目的波兰语言与文化能力基准测试,分为历史、地理、文化与传统、艺术与娱乐、语法、词汇六个类别。研究中,我们对超过30个开源与商业大模型进行了广泛评估,提供了超越传统自然语言处理任务与通用知识测评的新视角。

原文摘要 · Abstract (English)

Large language models (LLMs) are becoming increasingly proficient in processing and generating multilingual texts, which allows them to address real-world problems more effectively. However, language understanding is a far more complex issue that goes beyond simple text analysis. It requires familiarity with cultural context, including references to everyday life, historical events, traditions, folklore, literature, and pop culture. A lack of such knowledge can lead to misinterpretations and subtle, hard-to-detect errors. To examine language models' knowledge of the Polish cultural context, we introduce the Polish linguistic and cultural competency benchmark, consisting of 600 manually crafted questions. The benchmark is divided into six categories: history, geography, culture & tradition, art & entertainment, grammar, and vocabulary. As part of our study, we conduct an extensive evaluation involving over 30 open-weight and commercial LLMs. Our experiments provide a new perspective on Polish competencies in language models, moving past traditional natural language processing tasks and general knowledge assessment.

语言模型文化理解多语言评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。