arXiv:2608.05411cs.AI2026-08中稿 · and presented at t…

用新指标评估AI助教是否讲对了时机,还能提升教学效果。

Evaluating and Improving Pedagogical Fit in LLM-Based AI Tutors with the Pedagogical Suitability Index

  • 设计六维教学契合度评分,衡量AI回答是否匹配学生水平和课程进度。
  • 82.3%的差评案例经反馈后显著改善,证明可优化。
  • 适合教育科技研发者与在线教学平台,关注教学逻辑而非仅答案正确。

大型语言模型(LLMs)正被广泛用作AI助教,但正确答案未必是合适的教学回应。课堂中有效的帮助不仅需准确,还需符合学习者的当前基础、课程顺序及概念引入时机。现有评估多聚焦答案质量,忽视教学适配性。本文提出教学契合度指数(PSI),一个由六项理论驱动子分数组成的综合指标,用于评估LLM生成的辅导回应与学习者准备状态和课程进展的匹配程度,并将PSI作为结构化反馈信号以改进回答。在240个场景化测试中对比四种模型(ChatGPT、Gemini、Gemma4、Qwen3),使用标准与缺陷提示对,发现四者总体差距小(PSI范围:0.557至0.638),开源与闭源模型无明显区分。在提示扰动下,整体PSI基本稳定(Δ = -0.002),但子分项出现权衡。对62个低分案例应用PSI引导重生成,51例改善(82.3%)。人工评估确认弱项具有教学意义,多数改写获人类认可。结果表明,教学适配性可能比模型类型更关键,且可测量、可提升。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used as AI tutors, but a correct answer is not always a pedagogically appropriate one. In classroom learning, effective help depends not only on correctness, but also on whether a response matches the learner's current foundation, the course sequence, and the timing of concept introduction. Existing evaluations focus mainly on answer quality, leaving this instructional fit under-measured. We present the Pedagogical Suitability Index (PSI), a composite metric of six theory-informed sub-scores that evaluates how well LLM-generated tutoring responses align with learner readiness and curricular progression, and we further use PSI as a structured feedback signal for response improvement. We evaluate four LLM tutors (ChatGPT, Gemini, Gemma4, and Qwen3) across 240 scenario-based evaluations using paired standard and defective prompts, then apply a PSI-guided regeneration protocol to 62 weak-performing cases. Baseline differences across the four tested models were modest overall (PSI range: 0.557 to 0.638), and open-weight and closed models did not exhibit a clear separation in pedagogical fit. Under the tested prompt perturbations, overall PSI remained largely stable (Delta = -0.002), though sub-score trade-offs emerged. More importantly, PSI-guided feedback substantially improved weak-performing cases: 51 of 62 cases improved (82.3%). Focused manual evaluation of the 62 PSI-selected weak cases provides initial evidence that the identified weaknesses are instructionally meaningful and that many PSI-guided regenerations correspond to human-judged improvement. These results suggest that learner- and curriculum-aware alignment may matter more for effective tutoring than model category alone, and that such alignment is both measurable and improvable.

AI助教教学评估提示工程教育科技

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。