构建智能导师测评平台,评估AI作导师或学生的表现。
TutorGym: A Testbed for Evaluating AI Agents as Tutors and Students
- 将AI放入真实教学系统界面,交互式评测其教学与学习能力。
- 当前大模型做导师准确率仅52%-70%,随机水平以下;但作为学生能模拟人类学习曲线。
- 适合研究教育AI、认知建模及强化学习在教学中的应用者使用。
大型语言模型在MATH、GSM8K等学术基准上的进步,推动其作为独立导师或人类学习模拟器的应用。然而,这些应用需超越最终解法生成的评估。本文提出TutorGym,一个标准接口,用于在已通过课堂验证的智能辅导系统(如Cognitive Tutors、Apprentice Tutors、OATutors)中测试AI代理。TutorGym不仅提供问题-解法基准,更将AI置于现有教学系统的交互界面中,每一步均要求其以导师或学生身份决策。作为导师,模型需生成示例、提示和步骤级反馈,可直接与已有系统对比;作为学生,其学习过程与错误轨迹可与真实学生数据比对。TutorGym为训练与评估多种AI代理(包括LLMs、学习计算模型、强化学习代理)提供统一框架,涵盖223个不同教学领域。初步评估显示,当前大模型作导师表现不佳——无法有效识别错误操作,下一步动作正确率仅52%-70%;但作为学生,通过上下文学习可生成极类人的学习轨迹。
原文摘要 · Abstract (English)
Recent improvements in large language model (LLM) performance on academic benchmarks, such as MATH and GSM8K, have emboldened their use as standalone tutors and as simulations of human learning. However, these new applications require more than evaluations of final solution generation. We introduce TutorGym to evaluate these applications more directly. TutorGym is a standard interface for testing artificial intelligence (AI) agents within existing intelligent tutoring systems (ITS) that have been tested and refined in classroom studies, including Cognitive Tutors (CTAT), Apprentice Tutors, and OATutors. TutorGym is more than a simple problem-solution benchmark, it situates AI agents within the interactive interfaces of existing ITSs. At each step of problem-solving, AI agents are asked what they would do as a tutor or as a learner. As tutors, AI agents are prompted to provide tutoring support -- such as generating examples, hints, and step-level correctness feedback -- which can be evaluated directly against the adaptive step-by-step support provided by existing ITSs. As students, agents directly learn from ITS instruction, and their mistakes and learning trajectories can be compared to student data. TutorGym establishes a common framework for training and evaluating diverse AI agents, including LLMs, computational models of learning, and reinforcement learning agents, within a growing suite of learning environments. Currently, TutorGym includes 223 different tutor domains. In an initial evaluation, we find that current LLMs are poor at tutoring -- none did better than chance at labeling incorrect actions, and next-step actions were correct only ~52-70% of the time -- but they could produce remarkably human-like learning curves when trained as students with in-context learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。