用加纳科学数学竞赛题构建全球南方的推理评估基准
NSMQ Riddles: A Benchmark of Scientific and Mathematical Riddles for Quizzing Large Language Models
- 基于加纳全国科学数学竞赛的11年谜题,每题含至少3个线索
- 1800道题挑战顶尖大模型,表现仍不及优秀中学生
- 提供可自动评估的答案格式,适合测试全球教育场景下的推理能力
大型语言模型在多种科学教育基准上表现良好,显示出在科学与数学教育中的潜力。然而,现有评估多集中于西方数据集,缺乏全球南方代表性,且常采用易于判断的多项选择题。本文提出NSMQ Riddles,一个来自加纳国家科学与数学竞赛(NSMQ)的全新基准,涵盖11年5轮共1800道科学与数学谜题,每道题至少包含3个线索。参赛者需在提示逐步揭示时抢先作答,早期线索更模糊但得分更高。答案为数字、单词或短语,便于自动评估。我们评估了闭源(GPT-5.4、Gemini 3.1 Pro、Claude Opus 4.6)和开源(Kimi-K2.5、DeepSeek-V3.1、GPT-OSS-120B)先进模型,采用高/低推理设置。结果表明,该数据集对顶尖模型仍具挑战性,其表现低于最佳学生选手。本工作为全球南方科学数学推理能力提供了新颖且具有挑战性的评估基准,推动真正全球化的语言模型教育能力评测。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have shown good performance on various science educational benchmarks, demonstrating their potential for use in science and mathematics education. Yet, LLMs tend to be evaluated on science and mathematical educational datasets from the Western world, with an underrepresentation of datasets from the Global South. Furthermore, they tend to have multiple-choice answer options that are trivial to evaluate. In this work, we present NSMQ Riddles, a novel benchmark of Scientific and Mathematical Riddles from Ghana's National Science and Maths Quiz (NSMQ) competition to evaluate LLMs. The NSMQ is an annual live TV competition for senior secondary school students in Ghana that brings together the smartest high school students in Ghana who compete in teams of 2 by answering questions in biology, chemistry, physics, and math over five rounds and five stages until a winning team is crowned for that year. NSMQ Riddles consists of 11 years of riddle questions (n=1.8K) from the 5th round, with each riddle containing a minimum of 3 clues. Students compete to be the first to guess the answer on any of the clues, with earlier clues being vague and also fetching more points. The answers are usually a number, word, or short phrase, allowing for automatic evaluation. We evaluated state-of-the-art models: closed (GPT-5.4, Gemini 3.1 Pro, Claude Opus 4.6) and open models (Kimi-K2.5, DeepSeek-V3.1, GPT-OSS-120B) with high and low reasoning settings. Our evaluation shows that the dataset is challenging even for state-of-the-art LLMs, which performed worse than the best student contestants. This work contributes a novel and challenging benchmark for scientific and mathematical reasoning from the Global South towards enabling a true global benchmarking of LLMs' capabilities for science and mathematics education.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。