arXiv:2504.07749cs.CLcs.AI2025-04ACL被引 8

构建首个覆盖挪威语双标准的生成与理解评测基准

NorEval: A Norwegian Language Understanding and Generation Evaluation Benchmark

  • 24个高质量人工标注数据集,含5个全新创建
  • 涵盖理解与生成任务,建立双语(书面挪威语)人类基线
  • 开源完整评测框架,支持19个开源模型对比

本文提出NorEval,一个面向挪威语生成语言模型的大规模标准化评测基准。NorEval包含24个高质量人工创建的数据集,其中5个为全新构建。与现有挪威语评测基准不同,NorEval覆盖广泛的自然语言理解与生成任务,确立人类基线,并同时关注挪威语的两种官方书面形式:博克默尔语(Bokmål)和新挪威语(Nynorsk)。所有数据集及超过100条人工编写的提示已集成至语言模型评估工具包(LM Evaluation Harness),支持灵活且可复现的评估。本文详细阐述了NorEval的设计,并报告了在多种场景下对19个开源预训练及指令微调模型的评测结果。该基准、评估框架及标注材料均已公开。

原文摘要 · Abstract (English)

This paper introduces NorEval, a new and comprehensive evaluation suite for large-scale standardized benchmarking of Norwegian generative language models (LMs). NorEval consists of 24 high-quality human-created datasets -- of which five are created from scratch. In contrast to existing benchmarks for Norwegian, NorEval covers a broad spectrum of task categories targeting Norwegian language understanding and generation, establishes human baselines, and focuses on both of the official written standards of the Norwegian language: Bokmål and Nynorsk. All our datasets and a collection of over 100 human-written prompts are integrated into LM Evaluation Harness, ensuring flexible and reproducible evaluation. We describe the NorEval design and present the results of benchmarking 19 open-source pre-trained and instruction-tuned LMs for Norwegian in various scenarios. Our benchmark, evaluation framework, and annotation materials are publicly available.

语言模型挪威语评测基准生成任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。