用AI评估导师应对教育不公的能力,效果良好且成本可控。
Do Tutors Learn from Equity Training and Can Generative AI Assess It?
- 用GPT-4o等大模型分析导师应对不公情境的回应质量。
- 81名本科生导师自评信心提升,显示培训有微弱成效。
- 适合教育公平研究者、AI评估系统开发者参考。
教育公平是学习分析的核心议题,但缺乏能规模化开展的教学与评估工具,尤其受限于语言评价难题。本文探索大语言模型(LLMs)在教育公平领域的应用,评估81名远程本科生导师在在线课程中的表现。通过混合方法分析发现,导师在前后测中自评应对中学生潜在不公情境的信心略有提升。GPT-4o和GPT-4-turbo均表现出色,可准确预测并解释最优应对策略。综合性能、效率与成本,少样本学习下的GPT-4o为首选模型。本研究公开了课程日志数据、导师回复、人工标注评分标准及生成式AI提示模板。未来工作将统一情景难度,并优化提示以支持大规模评分。
原文摘要 · Abstract (English)
Equity is a core concern of learning analytics. However, applications that teach and assess equity skills, particularly at scale are lacking, often due to barriers in evaluating language. Advances in generative AI via large language models (LLMs) are being used in a wide range of applications, with this present work assessing its use in the equity domain. We evaluate tutor performance within an online lesson on enhancing tutors' skills when responding to students in potentially inequitable situations. We apply a mixed-method approach to analyze the performance of 81 undergraduate remote tutors. We find marginally significant learning gains with increases in tutors' self-reported confidence in their knowledge in responding to middle school students experiencing possible inequities from pretest to posttest. Both GPT-4o and GPT-4-turbo demonstrate proficiency in assessing tutors ability to predict and explain the best approach. Balancing performance, efficiency, and cost, we determine that few-shot learning using GPT-4o is the preferred model. This work makes available a dataset of lesson log data, tutor responses, rubrics for human annotation, and generative AI prompts. Future work involves leveling the difficulty among scenarios and enhancing LLM prompts for large-scale grading and assessment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。