评测大模型解释能力,发现其教学适配性不足。
ELI-Why: Evaluating the Pedagogical Utility of Language Model Explanations
- 构建13.4万条'为什么'问题数据集,评估模型解释的教育适配性。
- GPT-4解释匹配目标学段仅50%,比人工标注低29个百分点。
- 用户认为模型解释比人工解释更不贴合自身需求,平均差20%。
当前语言模型在教育领域广泛应用,但其针对不同知识背景学习者调整回答的能力仍待深入探索。为此,我们提出ELI-Why,一个包含13.4万条“为什么”问题的基准数据集,用于评估语言模型的教育辅助能力。通过两项大规模人工研究,我们评估了模型生成解释在小学、中学和研究生三个教育阶段的适用性。第一项研究中,人类评分者以教育者身份评估解释与学段的匹配度,发现GPT-4生成的解释仅50%符合目标学段,而普通人编写的解释达79%。第二项研究中,评分者以学习者身份判断解释是否满足自身信息需求,结果显示,无论学段如何,用户普遍认为GPT-4生成的解释比人工编写的平均差20%。此外,自动评估指标显示,不同模型家族生成的解释在学段级别上难以区分,限制了其教学有效性。
原文摘要 · Abstract (English)
Language models today are widely used in education, yet their ability to tailor responses for learners with varied informational needs and knowledge backgrounds remains under-explored. To this end, we introduce ELI-Why, a benchmark of 13.4K "Why" questions to evaluate the pedagogical capabilities of language models. We then conduct two extensive human studies to assess the utility of language model-generated explanatory answers (explanations) on our benchmark, tailored to three distinct educational grades: elementary, high-school and graduate school. In our first study, human raters assume the role of an "educator" to assess model explanations' fit to different educational grades. We find that GPT-4-generated explanations match their intended educational background only 50% of the time, compared to 79% for lay human-curated explanations. In our second study, human raters assume the role of a learner to assess if an explanation fits their own informational needs. Across all educational backgrounds, users deemed GPT-4-generated explanations 20% less suited on average to their informational needs, when compared to explanations curated by lay people. Additionally, automated evaluation metrics reveal that explanations generated across different language model families for different informational needs remain indistinguishable in their grade-level, limiting their pedagogical effectiveness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。