arXiv:2607.28128cs.CLcs.AI2026-07被引 1

用大模型评估教学效果时,通用好评标准不可靠,需搭配具体教学指标。

Rethinking LLM-Judged Helpfulness as a Pedagogy Signal: A Pre-Registered Audit Across Tutor Models

  • 同一模型下对比直接给答案与教学引导策略
  • 通用评分标准无法区分教学效果,但专门评估可完全区分
  • 适合教育类AI评测或教学型模型开发人员参考

大模型辅导存在测量难题:通用的有用性评分能否区分直接给答案与教学引导?我们通过预注册研究进行审计。在三个辅导模型基座中,使用相同基础模型和固定弱模拟学生,比较对话式与教学式策略。以冻结的Claude Opus 4.8为统一评判标准,其评分固定后,再用GPT-5.6 Sol对1,179个确认阶段的辅导回合进行事后稳健性检验。在主基座上,两种策略在有用性评分上无显著差异,但在教学性评分上完全分离(Cliff's $|δ|=0.10$ 对 $1.0$)。跨两名评判者,教学性差异方向一致,而有用性排序依赖评判者,在三个基座中有两个出现反转。在仅使用Opus的消融实验中,七个主基座策略在平均教学性上相差2.3分,而有用性均值仅在0.25分范围内波动。此外,透露答案的回合后,学生独立工作量均下降,该结果不依赖评判者。在控制环境下,通用有用性不是可靠的教学信号。辅导评估应结合教学针对性评分与确定性过程指标。

原文摘要 · Abstract (English)

LLM tutoring poses a measurement problem: can a general-purpose helpfulness rubric distinguish direct answer-giving from pedagogical guidance? We audit this signal in a pre-registered study. Within each of three tutor bases, we compare conversational and pedagogical policies instantiated with the same underlying model and paired with one fixed weak simulated student. Deterministic detectors measure answer leakage and next-turn independent work. Claude Opus 4.8 is the frozen, condition-blind primary judge. After the Opus scores were fixed, GPT-5.6 Sol was prospectively specified for a post hoc robustness audit of the same 1,179 confirmatory answer-phase tutor turns under the frozen helpfulness and pedagogy rubrics. On the primary base under Opus, the policies do not differ significantly in helpfulness but are perfectly rank-separated under the pedagogy rubric (Cliff's $|δ|{=}0.10$ vs. $1.0$). Across the two judges, pedagogy contrasts retain their direction where detected, whereas the helpfulness ordering is judge-contingent, reversing between judges on two of three bases. In an Opus-only ablation, seven primary-base policies span $2.3$ points in mean judged pedagogy within a $0.25$-point band of mean judged helpfulness. Separately, answer-revealing turns are followed by less independent student work on every base, a result that is judge-invariant by construction. In this controlled setting, general-purpose helpfulness is not a reliable pedagogy signal. Tutor evaluation should pair pedagogy-targeted rubrics with deterministic process measures.

大模型辅导教学评估模型评测人工评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。