用心理测量学方法构建教育领域大模型能力评估新基准。
A Novel Psychometrics-Based Approach to Developing Professional Competency Benchmark for Large Language Models
- 基于认知理论和专家协作设计评估框架。
- 在俄语GPT模型上测试,发现高阶认知任务表现不足。
- 适合关注AI教育应用可靠性的研究者与开发者。
大语言模型(LLM)时代不仅带来训练挑战,更需解决评估问题。尽管已有诸多基准,但缺乏科学、可靠的评估方法。本文采用证据中心设计(ECD)方法论,提出基于严格心理测量原则的基准开发新路径。首次以教育学领域为例,构建融合布卢姆分类学的新型基准,由受过测评开发训练的教育专家团队共同设计。该基准非针对人类,而是专为评估大模型而设。在俄语GPT模型上实证测试,结果显示模型在不同复杂度任务中表现不均,尤其在需要深层认知参与的任务上存在明显短板。研究表明,生成式AI虽具潜力,可用于个性化辅导、实时反馈和多语言学习,但当前作为独立教师助手的可靠性仍有限。
原文摘要 · Abstract (English)
The era of large language models (LLM) raises questions not only about how to train models, but also about how to evaluate them. Despite numerous existing benchmarks, insufficient attention is often given to creating assessments that test LLMs in a valid and reliable manner. To address this challenge, we accommodate the Evidence-centered design (ECD) methodology and propose a comprehensive approach to benchmark development based on rigorous psychometric principles. In this paper, we have made the first attempt to illustrate this approach by creating a new benchmark in the field of pedagogy and education, highlighting the limitations of existing benchmark development approach and taking into account the development of LLMs. We conclude that a new approach to benchmarking is required to match the growing complexity of AI applications in the educational context. We construct a novel benchmark guided by the Bloom's taxonomy and rigorously designed by a consortium of education experts trained in test development. Thus the current benchmark provides an academically robust and practical assessment tool tailored for LLMs, rather than human participants. Tested empirically on the GPT model in the Russian language, it evaluates model performance across varied task complexities, revealing critical gaps in current LLM capabilities. Our results indicate that while generative AI tools hold significant promise for education - potentially supporting tasks such as personalized tutoring, real-time feedback, and multilingual learning - their reliability as autonomous teachers' assistants right now remain rather limited, particularly in tasks requiring deeper cognitive engagement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。