arXiv:2608.07497cs.HCcs.CL2026-08

开源框架评估对话式学习模拟的真实性与教学互动质量。

EvalConvoLearn: An Open-Source Framework for Evaluating Grounded Learner Simulations in Tutoring Conversations

论文配图:EvalConvoLearn: An Open-Source Framework for Evaluating Grounded Learner Simulations in Tutoring Conversations
图 1 · 摘自论文原文
  • 基于真实对话数据,从学习表现和对话质量双维度评估模拟学习者。
  • 量化分析答案分布、提问频率、话轮长度等指标,贴近真实学生行为。
  • 适合教育AI研究者、智能导师系统开发者使用,支持代码复现。

对话式学习模拟是测试学习理论、评估教学材料与自动导师系统、或构建可教代理的重要工具。近年来,大语言模型(LLM)使模拟学习者的交互更加丰富自然;然而,尚无公开框架用于评估这些模拟是否忠实还原真实学习者行为。本文提出 EvalConvoLearn,一个开源评估框架,从两个维度衡量学习模拟:学习行为(基于技能的掌握结果)与对话质量(话轮策略、错误类型分布、提问率、话轮长度)。该框架通过将评估指标锚定在真实辅导对话数据集上,并以现有导师语句为基础生成回应,来衡量模拟学习者与真实数据中答案分布的接近程度。框架在一组辅导对话数据集上进行了演示,包含两个基于 LLM 的学习者模拟结果,以及已发布的 GitHub 代码。

原文摘要 · Abstract (English)

Conversational learner simulations are valuable tools for testing learning theories, evaluating instructional materials and automated tutors, or powering teachable agents. Recently, large language models (LLM) have enabled richer, more naturalistic interactions with simulated learners; however, no open framework exists for evaluating whether such simulations faithfully reproduce real learner behavior. We introduce EvalConvoLearn, an open-source framework that assesses learner simulations along two axes: learning behavior (skill-conditioned mastery outcomes) and conversational quality (talk moves, error type distributions, question rate, turn length). EvalConvoLearn measures how closely a simulated learner approximates answer distributions observed in data by grounding metrics in authentic tutoring conversation datasets, and anchoring generated tutor responses in existing tutor utterances. The framework is demonstrated on a dataset of tutoring dialogues, including results for two LLM-based learner simulations, and the published GitHub code.

教育AI对话评估模拟学习者大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。