首个评估虚拟学生主观能力边界的教育基准,助力打造更可信的智能教学助手。
EduPersona: Benchmarking Subjective Ability Boundaries of Virtual Student Agents
- 构建跨语言、多学科的虚拟学生对话数据集,涵盖10种人格类型
- 三阶段评估框架使虚拟学生表现提升30%以上,尤其在长期人格一致性上进步显著
- 适合教育AI研究者与开发者,推动更人性化的智能助教发展
随着大语言模型在教育领域的深入应用,虚拟学生代理在课堂模拟与教师培训中日益重要。然而其面向课堂的主观能力仍缺乏系统评估,限制了对模型能力边界的理解和可信部署。我们提出EduPersona,一个覆盖两种语言、三个学科、十种基于五大性格理论的人格类型的大型基准。数据集包含1,308轮真实课堂对话,对应12,814次师生问答,通过人格风格化扩展至约12.8万次交互,为评估提供坚实基础。在此基础上,我们将难以量化的主观表现分解为三个递进任务:任务1基础连贯性(行为、情绪、表达与语音是否符合课堂情境),任务2学生真实性,任务3长期人格一致性,建立基于教育理论的评估框架。我们在三个代表性LLM上进行系统实验,对比原始版本与在EduPersona上微调的十种人格变体。结果表明,在所有任务中均有显著且一致的提升:任务1提升33.6%,任务2提升30.6%,任务3提升14.9%。这验证了数据集的有效性与研究价值,同时揭示人格建模难度的异质性。EduPersona是首个聚焦主观能力的课堂基准,建立了可解耦、可验证的研究范式,我们将开源数据集与框架,支持社区推进可信、类人的教育AI发展。
原文摘要 · Abstract (English)
As large language models are increasingly integrated into education, virtual student agents are becoming vital for classroom simulation and teacher training. Yet their classroom-oriented subjective abilities remain largely unassessed, limiting understanding of model boundaries and hindering trustworthy deployment. We present EduPersona, a large-scale benchmark spanning two languages, three subjects, and ten persona types based on the Big Five theory. The dataset contains 1,308 authentic classroom dialogue rounds, corresponding to 12,814 teacher-student Q&A turns, and is further expanded through persona stylization into roughly 10 times larger scale (128k turns), providing a solid foundation for evaluation. Building on this resource, we decompose hard-to-quantify subjective performance into three progressive tasks: TASK1 basic coherence (whether behavior, emotion, expression, and voice align with classroom context), TASK2 student realism, and TASK3 long-term persona consistency, thereby establishing an evaluation framework grounded in educational theory and research value. We conduct systematic experiments on three representative LLMs, comparing their original versions with ten persona-fine-tuned variants trained on EduPersona. Results show consistent and significant average improvements across all tasks: TASK1 +33.6%, TASK2 +30.6%, and TASK3 +14.9%. These improvements highlight the dataset's effectiveness and research value, while also revealing the heterogeneous difficulty of persona modeling. In summary, EduPersona delivers the first classroom benchmark centered on subjective abilities, establishes a decoupled and verifiable research paradigm, and we will open-source both the dataset and the framework to support the broader research community in advancing trustworthy and human-like AI for education.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。