arXiv:2607.19209cs.CYcs.AI2026-07中稿 · the main track of …

用聚类和大模型评估编程团队协作,提升反馈速度与准确性

Assessment in Team Problem-Solving Exercises in Computing Education

论文配图:Assessment in Team Problem-Solving Exercises in Computing Education
图 1 · 摘自论文原文
  • 通过聚类分析团队行为模式,实现快速分组反馈
  • GPT-5.2相比GPT-4o更接近教师评分,误差显著降低
  • 方法已集成至开源平台INJECT,支持教学实践推广

本文针对计算教育中桌面推演(TTX)的团队协作评估难题,提出并比较了两种基于数据的评估方法:聚类分析与大语言模型(LLM)。研究基于来自两个国家共81名参与者的真实数据,将这些方法与教师依据标准化量规给出的评分进行对比。聚类方法能有效识别相似任务应对策略的团队,使教师可快速提供针对性反馈,具有高可靠性且计算成本低。而使用标准化量规的LLM评估中,GPT-4o频繁偏离人工评分,但GPT-5.2表现显著改善,误差大幅降低。所提方法已集成至开源的TTX学习平台INJECT,以支持教学规模化应用。为促进社区采用,研究公开所有数据集、工具及完整推演场景。

原文摘要 · Abstract (English)

This full paper in the research-to-practice track presents methods for assessing student teams in tabletop exercises (TTXs). TTXs enable learner teams to prepare for workplace tasks and practice crisis responses, such as resolving cybersecurity incidents. While assessment is essential for determining how well teams achieve learning objectives, the complex, open-ended nature of TTXs often leads to delayed or incomplete feedback. TTX learning platforms can record teams' actions and communication; yet, leveraging these data to assess performance is underexplored. To address this gap, we compared two post-TTX team assessment methods -- clustering and large language models (LLMs) -- using an original dataset from 81 participants across two countries. We evaluated these methods against instructor-assigned scores based on standardized rubrics. Clustering grouped teams that approached TTX tasks similarly, enabling instructors to deliver faster, targeted feedback to teams within a cluster. This method was valid and reliable, with low computational requirements. LLMs used the standardized rubrics to assess teams' communication. While GPT-4o frequently disagreed with instructor scores, GPT-5.2 demonstrated considerably lower error. The researched methods have been integrated into INJECT, an open-source TTX learning platform, to support scalability and teaching practice. To encourage community adoption, we publicly share all datasets, software tools, and a full-fledged TTX scenario.

团队评估大模型应用教育技术开源平台

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。