arXiv:2502.15127cs.AIcs.HC2025-02被引 1

用双阶段类图灵测试验证AI是否真懂学生思维。

The Imitation Game for Educational AI

  • 通过学生错题生成干扰项,检验AI能否模仿人类专家认知。
  • 若学生选择AI生成干扰项的比率接近人类专家,说明AI理解学生思维。
  • 适合教育AI研发者、智能辅导系统设计者使用。

随着人工智能在教育中的普及,一个核心挑战浮现:如何验证AI是否真正理解学生思考与推理方式?传统评估方法如学习成效测量需长期研究且受多种变量干扰。本文提出基于双阶段类图灵测试的新评估框架。第一阶段,学生对问题提供开放式回答,暴露自然错误概念;第二阶段,基于学生具体错误,由AI和人类专家分别生成相关问题的干扰项。通过分析学生选择AI生成干扰项的频率是否与人类专家生成的相似,可验证AI是否具备建模学生认知的能力。我们证明该评估必须基于个体响应——无条件方法仅能捕捉普遍误解。借助严格的统计抽样理论,我们确立了高置信度验证所需条件。本研究将条件化干扰项生成视为探测AI建模学生思维能力的关键指标,这一能力可支持个性化辅导、反馈与评估。

原文摘要 · Abstract (English)

As artificial intelligence systems become increasingly prevalent in education, a fundamental challenge emerges: how can we verify if an AI truly understands how students think and reason? Traditional evaluation methods like measuring learning gains require lengthy studies confounded by numerous variables. We present a novel evaluation framework based on a two-phase Turing-like test. In Phase 1, students provide open-ended responses to questions, revealing natural misconceptions. In Phase 2, both AI and human experts, conditioned on each student's specific mistakes, generate distractors for new related questions. By analyzing whether students select AI-generated distractors at rates similar to human expert-generated ones, we can validate if the AI models student cognition. We prove this evaluation must be conditioned on individual responses - unconditioned approaches merely target common misconceptions. Through rigorous statistical sampling theory, we establish precise requirements for high-confidence validation. Our research positions conditioned distractor generation as a probe into an AI system's fundamental ability to model student thinking - a capability that enables adapting tutoring, feedback, and assessments to each student's specific needs.

教育AI认知建模评估方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。