arXiv:2607.18529cs.HCcs.AI2026-07

用三个AI角色协作评估教学视频质量,更准且可解释。

EduPanel: A Three-Agent LLM Judge for Teaching Videos -- Reliability, Complementarity, and Human Trust Calibration

论文配图:EduPanel: A Three-Agent LLM Judge for Teaching Videos -- Reliability, Complementarity, and Human Trust Calibration
图 1 · 摘自论文原文
  • 分角色协同评估:三类专用AI分别处理不同教学维度。
  • 接近人类专家水平:专家评分误差从0.87降到0.73。
  • 适合教育评价辅助:可被专家识别并信任,不盲目接受结果。

教学视频正成为主流教育媒介,亟需可扩展的评价体系。现有自动评估方法未能充分应对教学质量依赖多模态证据且应针对目标学习者的特点。我们提出EduPanel,一种基于评分标准、面向学习者的LLM评估框架,通过专业化代理分解评估任务,生成可解释的教学质量分析。在专家研究、结构消融实验和学习者角色分析中,EduPanel表现可靠性与中位人类专家相当。专家评估显示其反馈使评分误差(MAE)从0.87降至0.73,且专家仍能有效识别不可靠输出(AUC = 0.77),而非盲信。结果表明,EduPanel适合作为教育评价的有力助手,而非人类专家的替代品。

原文摘要 · Abstract (English)

Teaching videos are becoming a major medium for education, creating a growing need for scalable evaluation of their pedagogical quality. Existing automatic judges do not fully address this setting because teaching quality depends on multimodal evidence and should be evaluated with respect to the intended learner rather than as a universal property. We present EduPanel, a rubric-grounded, learner-conditioned LLM judge that decomposes evaluation across specialized agents to produce interpretable assessments for different aspects of teaching quality. Across expert studies, architecture ablations, and learner-persona analyses, EduPanel achieves reliability comparable to a median human expert. In expert evaluation, its feedback improves scoring accuracy (MAE 0.87 to 0.73), while experts remain able to detect unreliable outputs (AUC = 0.77) instead of accepting them blindly. These results suggest that EduPanel can serve as effective assistants for educational evaluation rather than replacements for human experts.

教学评估多智能体LLM应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。