用大模型模拟有缺陷的学生,让教师练习教学应对。
Toward a Benchmark for Controllable Simulation of Imperfect Students with Large Language Models
- 通过提示控制模型保留或遗忘特定技能,实现可控学生模拟。
- 在数学任务中成功诱导出部分掌握的技能表现,可量化评估。
- 适合教师培训、教育模拟和大模型行为控制研究者。
教师教育需要针对具有明确优劣势和部分掌握能力的学习者进行刻意练习。大语言模型可通过模拟具备已知技能组件的学生来支持此类练习,使教师能够反复演练讲解、诊断与教学回应。但关键不在于最大化基准准确率或消除孤立事实,而在于控制模型行为以反映指定技能分布。本文探讨了提示驱动的语言模型是否可被引导以保留某些技能同时抑制其他技能。我们提出一个面向基准的框架:显式技能向量表示模拟学生,基于提示的控制设定保留与缺失的能力,通过技能匹配度指标、保留与遗忘对比及跨技能校准分析评估行为。结果表明,在结构化的数学情境中可诱导并测量选择性部分掌握,但可控程度依赖于模型本身。这些发现将可控学习者模拟定位为教师教育、教育模拟与语言模型控制交叉领域的独立研究问题。
原文摘要 · Abstract (English)
Teacher education requires deliberate practice with learners who exhibit identifiable strengths, weaknesses, and partial mastery. Large language models could support such practice by simulating students with known skill components, enabling teachers to rehearse explanations, diagnoses, and instructional responses. For this purpose, however, the central requirement is neither to maximize benchmark accuracy nor to suppress isolated facts, but to control model behavior so that it reflects a specified skill profile. This paper investigates whether prompted language models can be steered to retain some skills while suppressing others. We introduce a benchmark-oriented framework in which an explicit skill vector represents a simulated student, prompt-based control specifies retained and missing competencies, and behavior is evaluated using profile-alignment metrics, retained-versus-forgotten comparisons, and cross-skill calibration analyses. The results show that selective partial mastery can be induced and measured in a structured mathematics setting, although the degree of controllability remains model-dependent. These findings position controllable learner simulation as a distinct research problem at the intersection of teacher education, educational simulation, and language-model control.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。