为欧洲葡萄牙语打造首个语言维度评估基准,填补多语言大模型评测空白。
ALBA: A European Portuguese Benchmark for Evaluating Language and Linguistic Dimensions in Generative LLMs
- 从语言学角度构建8个维度的欧洲葡语评测集
- 专家手工标注+大模型评分实现可扩展评估
- 揭示模型在不同语言维度上的表现差异,适合语言研究者使用
随着大语言模型向多语言领域拓展,评估其在低资源语言中的表现愈发重要。欧洲葡萄牙语(pt-PT)尤其面临挑战,因现有训练数据和评测基准主要基于巴西葡萄牙语(pt-BR)。为此,我们提出ALBA,一个从零开始构建的语言学基础评测集,用于评估大模型在欧洲葡语中的语言相关任务能力,涵盖语言变体、文化语义、话语分析、文字游戏、语法、形态、词源学以及音系学共八个维度。ALBA由语言专家手工构建,并搭配大模型作为评判者框架,实现对pt-PT生成文本的可扩展评估。在多种模型上的实验显示,各语言维度间性能差异显著,凸显了全面、敏感于语言多样性的评测基准的重要性,有助于推动pt-PT相关工具的发展。
原文摘要 · Abstract (English)
As Large Language Models (LLMs) expand across multilingual domains, evaluating their performance in under-represented languages becomes increasingly important. European Portuguese (pt-PT) is particularly affected, as existing training data and benchmarks are mainly in Brazilian Portuguese (pt-BR). To address this, we introduce ALBA, a linguistically grounded benchmark designed from the ground up to assess LLM proficiency in linguistic-related tasks in pt-PT across eight linguistic dimensions, including Language Variety, Culture-bound Semantics, Discourse Analysis, Word Plays, Syntax, Morphology, Lexicology, and Phonetics and Phonology. ALBA is manually constructed by language experts and paired with an LLM-as-a-judge framework for scalable evaluation of pt-PT generated language. Experiments on a diverse set of models reveal performance variability across linguistic dimensions, highlighting the need for comprehensive, variety-sensitive benchmarks that support further development of tools in pt-PT.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。