用大模型模拟学生对话,发现现有方法效果有限。
Simulated Students in Tutoring Dialogues: Substance or Illusion?
- 提出学生模拟任务的正式定义与多维度评估指标
- 实测显示提示词法表现差,微调仍不理想
- 适合教育AI评估研究者参考,推动真实学生数据构建
大语言模型推动教育技术革新,但评估新方案需真实学生参与,成本高且难扩展。因此,众多基于大模型的辅导系统采用模拟学生进行训练与评估,通常通过简单提示实现。然而,极少工作关注或衡量模拟学生的质量。本文首次正式定义学生模拟任务,提出涵盖语言、行为与认知层面的评估指标,并在真实数学辅导对话数据集上对多种模拟方法进行基准测试。自动化与人工评估结果均表明,提示策略表现不佳;监督微调与偏好优化虽有提升,但仍有限,凸显该任务的挑战性,亟需后续研究突破。
原文摘要 · Abstract (English)
Advances in large language models (LLMs) enable many new innovations in education. However, evaluating the effectiveness of new technology requires real students, which is time-consuming and hard to scale up. Therefore, many recent works on LLM-powered tutoring solutions have used simulated students for both training and evaluation, often via simple prompting. Surprisingly, little work has been done to ensure or even measure the quality of simulated students. In this work, we formally define the student simulation task, propose a set of evaluation metrics that span linguistic, behavioral, and cognitive aspects, and benchmark a wide range of student simulation methods on these metrics. We experiment on a real-world math tutoring dialogue dataset, where both automated and human evaluation results show that prompting strategies for student simulation perform poorly; supervised fine-tuning and preference optimization yield much better but still limited performance, motivating future work on this challenging task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。