arXiv:2512.13330cs.CLcs.AI2025-12

首个专为芬兰语大模型设计的统一评测基准,覆盖多种任务类型。

FIN-bench-v2: A Unified and Robust Benchmark Suite for Evaluating Finnish Large Language Models

  • 整合芬兰语版主流基准与原版升级扩展,形成统一格式数据集。
  • 通过4项评估标准筛选出21个稳健任务,确保评测可靠性。
  • 适合芬兰语NLP研究者、模型开发者及多语言评测需求者使用。

我们提出 FIN-bench-v2,一个用于评估芬兰语大语言模型的统一基准套件。该套件整合了广泛使用的基准的芬兰语版本,以及原始 FIN-bench 的更新和扩展版本,形成单一、格式一致的数据集集合,涵盖多项选择和生成式任务,覆盖阅读理解、常识推理、情感分析、世界知识和对齐性。所有数据集均已转换为 HuggingFace Datasets 格式,包含闭合填空和多项选择提示形式,每项任务均有五个变体。对于机器翻译资源(如 GoldenSwag 和 XED),我们引入人工标注或审校。为筛选稳健任务,我们预训练一组 2.15B 参数的解码器模型,并利用其学习曲线计算单调性、信噪比、非随机性能及模型排序一致性,仅保留满足所有标准的任务。我们进一步评估若干更大的指令微调模型,以刻画不同任务与提示形式下的性能表现。所有数据集、提示及评估配置均通过我们在 Language Model Evaluation Harness 的分支公开:https://github.com/LumiOpen/lm-evaluation-harness。补充资源发布于独立仓库:https://github.com/TurkuNLP/FIN-bench-v2。

原文摘要 · Abstract (English)

We introduce FIN-bench-v2, a unified benchmark suite for evaluating large language models in Finnish. FIN-bench-v2 consolidates Finnish versions of widely used benchmarks together with an updated and expanded version of the original FIN-bench into a single, consistently formatted collection, covering multiple-choice and generative tasks across reading comprehension, commonsense reasoning, sentiment analysis, world knowledge, and alignment. All datasets are converted to HuggingFace Datasets, which include both cloze and multiple-choice prompt formulations with five variants per task, and we incorporate human annotation or review for machine-translated resources such as GoldenSwag and XED. To select robust tasks, we pretrain a set of 2.15B-parameter decoder-only models and use their learning curves to compute monotonicity, signal-to-noise, non-random performance, and model ordering consistency, retaining only tasks that satisfy all criteria. We further evaluate a set of larger instruction-tuned models to characterize performance across tasks and prompt formulations. All datasets, prompts, and evaluation configurations are publicly available via our fork of the Language Model Evaluation Harness at https://github.com/LumiOpen/lm-evaluation-harness. Supplementary resources are released in a separate repository at https://github.com/TurkuNLP/FIN-bench-v2.

芬兰语大模型评测基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。