arXiv:2507.12674cs.CYcs.AI2025-07中稿 · the Main Conferenc…

用模拟学生编程修改来评估AI导师反馈效果,更贴近真实学习行为。

ParaStudent: Closing the Sim2Real Gap in User Simulators for AI Tutor Evaluation

  • 基于微调框架模拟新手编程改写,匹配真实学生代码分布。
  • 在反馈相关性和采纳成功率上达到0.80 AUC,优于基线模型。
  • 适合用于部署前的AI导师反馈预评估,提升测试真实性。

在部署前评估人工智能(AI)导师反馈效果,需预测学生参与度,通常依赖真实交互数据。本文提出ParaStudent,一种用于模拟新手编程修改的微调框架,以支持AI导师评估。相较于提示基线,ParaStudent生成的代码修改在功能、风格和语义上更接近真实学生代码分布。最佳变体在区分真实参与度高于或低于中位数的样本时,反馈相关性与成功采纳率的AUC均达0.80,而提示基线在成功采纳率上仅接近随机水平。结果表明,模拟参与度可有效支撑部署前反馈筛选。

原文摘要 · Abstract (English)

Evaluating Artificial Intelligence (AI) tutor feedback before deployment requires anticipating student engagement, typically assessed through real interaction data. We introduce ParaStudent, a fine-tuning framework for simulating novice programming revisions to support AI tutor evaluation. Compared with prompted baselines, ParaStudent's revisions more closely match real student code distributions across functional, stylistic, and semantic metrics. Our best variant achieves AUCs of 0.80 for both feedback relevance and successful uptake when distinguishing streams with real engagement above versus at or below the median, while prompted baselines remain near chance on successful uptake. These findings demonstrate the promise of simulated engagement for pre-deployment feedback triage.

AI导师代码生成仿真评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。