arXiv:2501.11128cs.CLcs.AI2025-01中稿 · NoDaLiDa / Baltic-…被引 5

构建了覆盖挪威语多种能力的问答数据集,助力本地化语言模型研究。

A Collection of Question Answering Datasets for Norwegian

  • 基于挪威语两种书写标准,由母语者构建超1万条问答对。
  • 多数模型在书面语(Bokmål)表现优于新挪威语(Nynorsk),常识推理最弱。
  • 适合研究挪威语自然语言理解、低资源语言模型及真相生成的学者。

本文介绍一套面向挪威语的新问答数据集,包括NorOpenBookQA、NorCommonSenseQA、NorTruthfulQA和NRK-Quiz-QA。数据涵盖世界知识、常识推理、真实性及挪威相关知识等多个领域,覆盖挪威语的两种书面形式——博克莫尔语(Bokmål)与新挪威语(Nynorsk),共包含超过10,000个问答对,均由母语者创建。我们详细阐述数据构建方法,并在零样本与少样本条件下评估了11个语言模型的表现。结果显示,多数模型在博克莫尔语中表现优于新挪威语,尤其在常识推理任务上表现最差,且常生成不真实答案。所有数据集及标注材料均公开可获取。

原文摘要 · Abstract (English)

This paper introduces a new suite of question answering datasets for Norwegian; NorOpenBookQA, NorCommonSenseQA, NorTruthfulQA, and NRK-Quiz-QA. The data covers a wide range of skills and knowledge domains, including world knowledge, commonsense reasoning, truthfulness, and knowledge about Norway. Covering both of the written standards of Norwegian - Bokmål and Nynorsk - our datasets comprise over 10k question-answer pairs, created by native speakers. We detail our dataset creation approach and present the results of evaluating 11 language models (LMs) in zero- and few-shot regimes. Most LMs perform better in Bokmål than Nynorsk, struggle most with commonsense reasoning, and are often untruthful in generating answers to questions. All our datasets and annotation materials are publicly available.

问答数据集挪威语语言模型常识推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。