arXiv:2608.03206cs.CYcs.AI2026-08

构建30天持续教学评估基准,验证AI导师真实育人效果

EduClaw-Bench: A Long-Horizon Benchmark for Pedagogical LLM Agents with Simulated Learners

论文配图:EduClaw-Bench: A Long-Horizon Benchmark for Pedagogical LLM Agents with Simulated Learners
图 1 · 摘自论文原文
  • 用知识追踪模型模拟学习者,实现30天连续教学评测
  • 10个代理适配器中无一能全程保持优质教学表现
  • 引入双维度评价体系,适合教育AI研发与评估者参考

大型语言模型(LLMs)在辅导、作文评分等教育应用中发挥重要作用,但现有系统多为单一任务解决方案。近期虽有将这些功能集成到学习管理系统(LMS)中的智能体,但教学具有长周期特性——学生需数日乃至数周逐步提升。目前尚无基准可评估智能体导师在长期关系中的表现。本文提出EduClaw-Bench,一个为期30天的持续教学评估基准,其模拟学习者基于知识追踪(KT)模型构建,该模型由真实学生数据训练,驱动其回答并反映知识掌握情况。共设计55种教学场景,对每个代理从学习成效、响应性、助益性三个核心维度,以及加涅和罗森希尔教学原则两个课程设计维度进行评分。助益性和课程维度由三名跨家族大模型评委共同判断。在三个基础模型层级上评估10个代理适配器,发现:教学品质取决于基础模型与代理协同作用,而非单独决定;几乎所有组合均无法在整个周期内维持良好表现。通过校准检查(ECE=0.049)和真实课堂实地研究确认,模拟学习者及其测量指标有效贴近现实。本工作为未来可信教育AI导师的发展迈出关键一步。

原文摘要 · Abstract (English)

Large language models (LLMs) power educational applications from tutoring to essay scoring, but each is a point solution to a single task, and only recently have these point solutions been integrated into agents operating over a learning management system (LMS). Yet tutoring is long-horizon, since a learner improves over days and weeks rather than in a single turn, and no benchmark evaluates an agent tutor across a sustained relationship. We introduce EduClaw-Bench, a benchmark that places an agent tutor in a continuous 30-day relationship with a simulated learner grounded in knowledge tracing (KT), whose knowledge-concept mastery, from a KT model trained on real-student data, drives its answers and is probed for learning gain across 55 scenarios. Each agent is scored on three primary axes (learning gain, responsiveness, and helpfulness) and two curriculum-design axes (Gagné and Rosenshine), with helpfulness and the curriculum axes judged by a cross-family panel of three LLM judges. Evaluating 10 agent adapters over three base-model tiers yields two findings that single-tier, single-session evaluation cannot reach. First, tutoring quality belongs to the base model and the agent harness together rather than either alone. Second, almost no combination sustains good tutoring over the full horizon. A calibration check ($\text{ECE}=0.049$) and a live-classroom field study confirm that the simulated learner and its measurements track reality. Our work is a step toward trustworthy AI tutors for future education.

教育AI智能导师长程评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。