通过行为数据与注意力检测,发现学生越频繁接受代码补全,反思能力越弱。
To Tab or Not to Tab: Measuring Critical Engagement in AI Code Completion Tools Using Behavioral Signals and Attention Checks

- 记录学生与代码补全的交互行为,结合注意力检测评估反思性参与度。
- 接受补全次数越多,注意力测试表现越差;停留时间越长,表现越好。
- 为提升编程反思能力提供可量化的数据支持,适合教育研究者使用。
AI代码补全工具(如Github Copilot)为学生编写程序提供代码建议,但近期定性研究指出学生未能批判性评估这些建议。我们提出Clover工具,记录学生与代码建议的交互行为,并引入注意力检查以探测编程过程中的反思性参与。基于文献构建了人工智能辅助编程的行为交互指标分类体系。分析显示:接受补全频率越高,注意力检查表现越低;停留时间越长,注意力检查表现越高。研究结论表明,编程过程数据与注意力检查可共同支持反思性学习,适用于教育技术与人机协作研究。
原文摘要 · Abstract (English)
AI code completion tools, such as Github Copilot, provide students with code suggestions to help them write programs. However, recent qualitative studies suggest that students fail to critically evaluate these suggestions. We present Clover, a code completion tool that logs students' interactions with code suggestions and additionally offers attention checks to probe reflective engagement during programming tasks. We also develop a taxonomy of behavioral interaction metrics for AI-assisted programming, informed by literature. We analyzed relationships between interaction patterns, engagement with attention checks, and task performance. We observed that higher rates of tab accept were associated with lower attention check performance, while increased dwell time was associated with higher attention check performance. We conclude by discussing how programming process data and attention checks might support reflective engagement in AI-assisted programming.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。