arXiv:2606.16206cs.AIcs.CL2026-06中稿 · EMNLP被引 2

评测大模型教学效果,不能只看解题对不对,还得看会不会教。

Beyond Helpfulness: A Teaching-over-Solving Diagnostic for Measuring Educational Impact in LLM Tutors

  • 通过解题与教学两类指标的差距,诊断模型真实教学能力。
  • 8个模型中解题与教学得分相关性仅0.421,排名常大幅变化。
  • 强调引导提问、提示分层等教学行为,更适合评估教育价值。

大语言模型被越来越多地用作教育辅导工具,但解题能力强并不等于教学支持有效。我们研究公开的LLM辅导基准是否能区分真正的学习支持与单纯输出答案的行为。提出一种轻量级诊断方法,基于解题导向与教学导向基准表现的差距。利用MathTutorBench公开榜单结果,发现这两类维度仅部分一致:在8个公开模型中,解题与教学综合得分相关性为0.421,多个模型在从解题转向教学评价时排名发生显著变化。进一步分析TutorBench样本显示,代理行为(如引导性提问、分层提示、非暴露式支架)在评分标准中明确体现,尤其在主动学习场景下更为突出。结果表明,教育影响评估不应将任务成功视为学习支持的充分代理。我们主张公开辅导基准应分别报告解题与教学得分,并更明确地披露以学生主体性为核心的评判标准。

原文摘要 · Abstract (English)

Large language models are increasingly proposed as educational tutors, yet stronger task-solving ability does not necessarily imply stronger learning support. Motivated by recent calls to measure the social impact of NLP systems in practice, we study whether public LLM tutoring benchmarks distinguish learning-supportive behavior from mere answer production. We propose a lightweight diagnostic based on the gap between solving-oriented and pedagogy-oriented benchmark performance. Using public MathTutorBench leaderboard results, we show that these dimensions are only partially aligned: across eight publicly reported models, the correlation between solving and pedagogy composites is 0.421, and several models shift meaningfully in rank when evaluation moves from solving to pedagogy. We then analyze the public TutorBench sample and show that agency-relevant behaviors are explicitly encoded in benchmark rubrics, especially in active-learning settings that reward guiding questions, calibrated hints, and non-disclosive scaffolding. Together, these findings suggest that educational-impact evaluation should not treat task success as a sufficient proxy for learning support. We argue that public tutoring benchmarks can better support positive-impact evaluation by reporting solving-oriented and pedagogy-oriented scores separately and by making disclosure-sensitive, student-agency-preserving criteria more explicit.

大模型教学教育评估教学能力评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。