首个评估大模型适龄答题能力的基准数据集,专为中小学教育设计。
EduAdapt: A Question Answer Benchmark Dataset for Evaluating Grade-Level Adaptability in LLMs
- 构建涵盖1-12年级、9个科学学科的近4.8万条带年级标签的问答对
- 大模型在低年级(1-5年级)回答准确率仍显著偏低
- 适合教育AI研发者、教学工具开发者及评测人员使用
大型语言模型正在改变教育,能回答问题、解释复杂概念并生成内容。然而,尽管在学术基准上表现良好,它们往往无法根据学生年级调整回答难度。这在K-12教育中尤为关键,因为年龄适宜的词汇和解释对学习效果至关重要。现有模型常为年幼学习者提供过深或模糊的回答,且缺乏标准化评估框架。为此,我们提出EduAdapt,一个包含近48,000条带年级标签的问答对的数据集,覆盖九个科学主题,涵盖1-12年级,分为四个年级组。我们评估了多种开源大模型,发现虽然模型越大表现越好,但在低年级(1-5年级)仍存在明显适应困难。本研究首次提供评估大模型适龄响应能力的数据集与评测框架,旨在通过更优训练与提示策略推动更具发展适配性的教育人工智能系统。代码与数据已公开于https://github.com/NaumanNaeem/EduAdapt。
原文摘要 · Abstract (English)
Large language models (LLMs) are transforming education by answering questions, explaining complex concepts, and generating content across a wide range of subjects. Despite strong performance on academic benchmarks, they often fail to tailor responses to students' grade levels. This is a critical need in K-12 education, where age-appropriate vocabulary and explanation are essential for effective learning. Existing models frequently produce outputs that are too advanced or vague for younger learners, and there are no standardized benchmarks to evaluate their ability to adjust across cognitive and developmental stages. To address this gap, we introduce EduAdapt, a benchmark of nearly 48k grade-labeled QA pairs across nine science subjects, spanning Grades 1-12 and grouped into four grade levels. We evaluate a diverse set of open-source LLMs on EduAdapt and find that while larger models generally perform better, they still struggle with generating suitable responses for early-grade students (Grades 1-5). Our work presents the first dataset and evaluation framework for assessing grade-level adaptability in LLMs, aiming to foster more developmentally aligned educational AI systems through better training and prompting strategies. EduAdapt code and datasets are publicly available at https://github.com/NaumanNaeem/EduAdapt.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。