arXiv:2410.22839cs.CLcs.AI2024-10中稿 · NoDaLiDa/Baltic-HL…被引 2

评测大模型对丹麦语与文化的理解能力,发现其表现有稳定的核心因素。

Danoliteracy of Generative Large Language Models

  • 构建8类场景的丹麦语能力评测基准,涵盖公民考试与社交媒体问答。
  • 模型表现与人类评价相关性达0.8,GPT-4和Claude Opus排名最高。
  • 95%性能差异由模型语言适配一致性决定,体现通用适应力因子。

生成式大语言模型(GLLMs)的技术热潮不仅限于英语,也推动了低资源语言如丹麦语的应用、投资与关注度提升。然而,由于缺乏适用的评估语料,这些模型在丹麦语上的能力长期难以量化验证。本文提出一项针对丹麦语文化理解能力(Danoliteracy)的GLLM评测基准,涵盖丹麦公民测试、抽象社交媒体问答等八类多样化场景。该小型基准可生成稳健排名,与人类反馈的相关性达ρ≈0.8;其中GPT-4和Claude Opus表现最优。分析显示,模型在各场景中的性能差异中,有95%可由一个核心因素解释,表明大模型在丹麦语上存在语言适应一致性的“g因子”。

原文摘要 · Abstract (English)

The language technology moonshot moment of Generative Large Language Models (GLLMs) was not limited to English: These models brought a surge of technological applications, investments, and hype to low-resource languages as well. However, the capabilities of these models in languages such as Danish were, until recently, difficult to verify beyond qualitative demonstrations due to a lack of applicable evaluation corpora. We present a GLLM benchmark to evaluate \emph{Danoliteracy}, a measure of Danish language and cultural competency across eight diverse scenarios such as Danish citizenship tests and abstractive social media question answering. This limited-size benchmark was found to produce a robust ranking that correlates to human feedback at $ρ\sim 0.8$ with GPT-4 and Claude Opus models achieving the highest rankings. Analyzing these model results across scenarios, we find one strong underlying factor explaining $95\%$ of scenario performance variance for GLLMs in Danish, suggesting a $g$ factor of model consistency in language adaptation.

大模型评测多语言丹麦语语言适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。