用语言水平考试测试大模型对卢森堡语的支持能力。
Testing Low-Resource Language Support in LLMs Using Language Proficiency Exams: the Case of Luxembourgish
- 用语言水平考试评估大模型对低资源语言的处理能力。
- 大型模型如Claude在卢森堡语考试中得分高,小型模型表现差。
- 考试成绩可预测模型在其他卢森堡语任务中的表现。
大型语言模型(LLMs)已成为科研与社会应用的重要工具。尽管全球广泛使用,但这些模型主要针对英语用户设计,在英语及其他主流语言上表现良好,而卢森堡语等低资源语言则被忽视,缺乏相应的评估工具与数据集。本研究探索语言水平考试作为卢森堡语评估工具的可行性。结果表明,大型模型(如Claude、DeepSeek-R1)通常得分较高,小型模型表现较弱。此外,语言考试成绩可有效预测模型在其他卢森堡语自然语言处理任务中的表现。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have become an increasingly important tool in research and society at large. While LLMs are regularly used all over the world by experts and lay-people alike, they are predominantly developed with English-speaking users in mind, performing well in English and other wide-spread languages while less-resourced languages such as Luxembourgish are seen as a lower priority. This lack of attention is also reflected in the sparsity of available evaluation tools and datasets. In this study, we investigate the viability of language proficiency exams as such evaluation tools for the Luxembourgish language. We find that large models such as Claude and DeepSeek-R1 typically achieve high scores, while smaller models show weak performances. We also find that the performances in such language exams can be used to predict performances in other NLP tasks in Luxembourgish.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。