首个面向爱沙尼亚语的大语言模型基准测试,评估多类能力。
Estonian Native Large Language Model Benchmark
- 基于7个本土语料构建评测集,覆盖语法、知识、摘要等任务。
- 对比6个基础模型与26个指令微调模型,发现商业模型表现更优。
- 使用人类评分与大模型评阅双验证,结果一致性强。
爱沙尼亚语的大语言模型评测基准尚不完善,缺乏对不同模型在爱沙尼亚语任务上性能的全面评估。本文提出一个新的爱沙尼亚语大语言模型基准,基于七个多样化的数据集,涵盖通用知识、领域知识、爱沙尼亚语语法与词汇理解、摘要能力、上下文理解等维度。所有数据均源自本地爱沙尼亚语资源,未使用机器翻译生成。我们对比了6个基础模型和26个指令微调的开源及商用模型。评估采用人工评价与大模型作为评判者(LLM-as-a-judge)两种方法。人工评分与基准评估结果呈现中到高度相关性,具体因数据集而异。以Claude 3.7 Sonnet作为评判者时,其评分与人工评分高度一致,表明顶尖大模型可有效辅助爱沙尼亚语模型的评估。
原文摘要 · Abstract (English)
The availability of LLM benchmarks for the Estonian language is limited, and a comprehensive evaluation comparing the performance of different LLMs on Estonian tasks has yet to be conducted. We introduce a new benchmark for evaluating LLMs in Estonian, based on seven diverse datasets. These datasets assess general and domain-specific knowledge, understanding of Estonian grammar and vocabulary, summarization abilities, contextual comprehension, and more. The datasets are all generated from native Estonian sources without using machine translation. We compare the performance of base models, instruction-tuned open-source models, and commercial models. Our evaluation includes 6 base models and 26 instruction-tuned models. To assess the results, we employ both human evaluation and LLM-as-a-judge methods. Human evaluation scores showed moderate to high correlation with benchmark evaluations, depending on the dataset. Claude 3.7 Sonnet, used as an LLM judge, demonstrated strong alignment with human ratings, indicating that top-performing LLMs can effectively support the evaluation of Estonian-language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。