为生成式AI教学系统设计基于教育学的评估框架
Pedagogy-driven Evaluation of Generative AI-powered Intelligent Tutoring Systems
- 提出以学习科学为基础的评估方法,强调教育理论指导
- 指出现有评估依赖主观标准,缺乏统一性和可比性
- 适合教育科技研究者与AI教育产品开发者参考
人工智能在教育领域的研究已有长期积累,智能辅导系统(ITS)通过融合技术进步、教育理论和认知心理学不断发展。生成式AI模型的突破加速了大语言模型(LLM)驱动的ITS发展,使其具备模拟人性化、教学丰富且认知挑战性强的辅导潜力。然而,由于缺乏可靠、普遍接受且以教育学为导向的评估框架与基准,这些系统的进展与影响仍难以追踪。当前多数基于对话的教育类ITS评估依赖主观协议和非标准化基准,导致结果不一致且泛化能力有限。本文回归主流开发思路,总结当前最先进的评估实践,通过精心设计的真实案例研究揭示关键挑战。基于前期跨学科AIED研究洞察,提出三个切实可行、理论扎实的研究方向,根植于学习科学原则,旨在建立公平、统一且可扩展的智能辅导系统评估方法。
原文摘要 · Abstract (English)
The interdisciplinary research domain of Artificial Intelligence in Education (AIED) has a long history of developing Intelligent Tutoring Systems (ITSs) by integrating insights from technological advancements, educational theories, and cognitive psychology. The remarkable success of generative AI (GenAI) models has accelerated the development of large language model (LLM)-powered ITSs, which have potential to imitate human-like, pedagogically rich, and cognitively demanding tutoring. However, the progress and impact of these systems remain largely untraceable due to the absence of reliable, universally accepted, and pedagogy-driven evaluation frameworks and benchmarks. Most existing educational dialogue-based ITS evaluations rely on subjective protocols and non-standardized benchmarks, leading to inconsistencies and limited generalizability. In this work, we take a step back from mainstream ITS development and provide comprehensive state-of-the-art evaluation practices, highlighting associated challenges through real-world case studies from careful and caring AIED research. Finally, building on insights from previous interdisciplinary AIED research, we propose three practical, feasible, and theoretically grounded research directions, rooted in learning science principles and aimed at establishing fair, unified, and scalable evaluation methodologies for ITSs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。