构建挪威语新闻摘要数据集,用于评估生成式模型的摘要能力。
Benchmarking Abstractive Summarisation: A Dataset of Human-authored Summaries of Norwegian News Articles
- 收集挪威语新闻及母语者撰写的三份高质量摘要
- 覆盖两种挪威书面语变体,提供双语对照
- 可用于评测模型与人类摘要的差距,适合语言模型研究者
我们引入了一个高质量的挪威语新闻文章人工摘要数据集,旨在评估生成式语言模型的抽象摘要能力。数据集中每篇文档均配有三位母语挪威语使用者撰写的三份候选黄金标准摘要,所有摘要均以挪威书面语的两种变体——博克马尔语(Bokmål)和新挪威语(Nynorsk)呈现。本文详细描述了数据构建过程,并对现有开源大语言模型在该数据集上的表现进行了评估。此外,还提供了人工手动评估结果,对比了人工摘要与模型生成摘要的质量差异。结果表明,该数据集为挪威语摘要任务提供了具有挑战性的基准测试环境。
原文摘要 · Abstract (English)
We introduce a dataset of high-quality human-authored summaries of news articles in Norwegian. The dataset is intended for benchmarking the abstractive summarisation capabilities of generative language models. Each document in the dataset is provided with three different candidate gold-standard summaries written by native Norwegian speakers, and all summaries are provided in both of the written variants of Norwegian -- Bokmål and Nynorsk. The paper describes details on the data creation effort as well as an evaluation of existing open LLMs for Norwegian on the dataset. We also provide insights from a manual human evaluation, comparing human-authored to model-generated summaries. Our results indicate that the dataset provides a challenging LLM benchmark for Norwegian summarisation capabilities
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。