arXiv:2509.04111cs.CL2025-09被引 7

构建跨306语言的阅读理解数据集,评估多语言模型真实理解能力。

MultiWikiQA: A Reading Comprehension Benchmark in 300+ Languages

  • 用大模型生成带原文答案的问题,再重述问题防简单匹配。
  • 覆盖122万样本,30种语言人工评测显示问题自然流畅。
  • 验证不同语言模型在多语言任务中的巨大性能差距。

我们提出一个新的阅读理解数据集 MultiWikiQA,涵盖306种语言,共包含1,220,757个样本。数据基于维基百科文章构建,利用大语言模型生成与文章内容相关的问答对,确保答案在原文中完全出现。随后对问题进行重述,以降低简单词汇匹配方法的有效性。我们对其中30种语言(含低资源与高资源语言)进行了众包人工评估,共156名参与者,所有语言的平均流畅度评分均高于“基本自然”,表明样本质量良好。我们评估了6种不同规模的编码器与解码器语言模型,结果表明该基准具有足够难度,且各语言间性能差异显著。数据集与评估结果均已公开。

原文摘要 · Abstract (English)

We introduce a new reading comprehension dataset, dubbed MultiWikiQA, which covers 306 languages and has 1,220,757 samples in total. We start with Wikipedia articles, which also provide the context for the dataset samples, and use an LLM to generate question/answer pairs related to the Wikipedia article, ensuring that the answer appears verbatim within the article. Next, the question is then rephrased to hinder simple word matching methods from performing well on the dataset. We conduct a crowdsourced human evaluation of the fluency of the generated questions, which included 156 respondents across 30 of the languages (both low- and high-resource). All 30 languages received a mean fluency rating above ``mostly natural'', showing that the samples are of good quality. We evaluate 6 different language models, both decoder and encoder models of varying sizes, showing that the benchmark is sufficiently difficult and that there is a large performance discrepancy amongst the languages. Both the dataset and survey evaluations are publicly available.

阅读理解多语言数据集大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。