发现学生用AI助教主要为抄答案,而非深度学习。
Your Students Don't Use LLMs Like You Wish They Did

- 提出六种自动化评估教育对齐度的计算指标
- 12,650条对话显示学生多在截止日前突击提问
- 适合教育AI研究者验证教学目标达成度
教育类NLP系统通常以参与度和满意度调查评估,但这些仅是教学目标的间接指标。本文引入六种计算指标,用于自动评估学生与AI对话的教育对齐程度。基于来自四门课程共500次对话、12,650条消息的分析,发现根本性错位:教师设计对话式辅导旨在促进持续学习,但学生主要将其用于获取答案。部署上下文是使用模式最强预测因子,超过学生偏好或系统设计:当AI工具可选时,使用集中在截止日前;当嵌入课程结构中,学生直接提问作业原题。全对话评估会忽略这些逐轮模式。本研究提出的指标将帮助教育对话系统研究者衡量是否实现教学目标。
原文摘要 · Abstract (English)
Educational NLP systems are typically evaluated using engagement metrics and satisfaction surveys, which are at best a proxy for meeting pedagogical goals. We introduce six computational metrics for automated evaluation of pedagogical alignment in student-AI dialogue. We validate our metrics through analysis of 12,650 messages across 500 conversations from four courses. Using our metrics, we identify a fundamental misalignment: educators design conversational tutors for sustained learning dialogue, but students mainly use them for answer-extraction. Deployment context is the strongest predictor of usage patterns, outweighing student preference or system design: when AI tools are optional, usage concentrates around deadlines; when integrated into course structure, students ask for solutions to verbatim assignment questions. Whole-dialogue evaluation misses these turn-by-turn patterns. Our metrics will enable researchers building educational dialogue systems to measure whether they are achieving their pedagogical goals.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。