让大模型不仅模仿学生行为,还模拟其真实思维过程。
INSIDE the Student's Mind: Jointly Modeling Latent Reasoning and Action in LLM Student Simulators

- 基于布卢姆认知分类构建多维度内部对话生成机制
- 在代码生成任务中匹配真实学生行为,推理对齐达57.9%最高
- 适合教育智能系统评估与个性化教学研究者使用
基于大语言模型的学生模拟器常只能复现表面行为,却难以捕捉其背后的真实推理。在教育领域,这种差距尤为明显——两名学生可能提交完全相同的答案,但动机迥异。本文提出内部学生对话(INSIDE)框架,通过联合微调大模型的行为与思考轨迹,使其不仅像学生一样行动,更像学生一样思考。INSIDE生成基于布卢姆认知分类的认知、情感与行为三维内部对话,并在配对的思考痕迹与行为数据上进行微调。我们在两个维度评估:模拟行为的真实性与生成内部对话的质量。结果表明,INSIDE在行为保真度和推理一致性上均有显著提升,代码生成匹配真实学生表现,推理对齐度达到各模型中的最高值57.9%。
原文摘要 · Abstract (English)
Large Language Model (LLM)-based simulators often reproduce observable actions but fail to capture the underlying reasoning behind them. In education, where student simulation is increasingly used for various applications such as evaluating tutoring systems, this gap is especially pronounced. Two students may submit identical submissions for entirely different reasons. We present INTERNAL STUDENT DIALOGUE (INSIDE), a student modeling framework that fine-tunes LLMs not only to act like students but also to think like them. INSIDE generates internal dialogue grounded in Bloom's Taxonomy across cognitive, affective, and action dimensions, and fine-tunes models on paired think traces and actions. We baseline against different prompting frameworks and evaluate on two axes: fidelity of simulated actions and quality of generated internal dialogue. Our evaluations show that INSIDE improves simulation fidelity in both action fidelity, matching code generation of real students, and reasoning alignment, achieving the highest alignment across models up to 57.9%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。