对比教师与AI评估设计思维作品,发现两者在创意判断上差异大。
Human or AI? Comparing Design Thinking Assessments by Teaching Assistants and Bots
- 用AI和教师分别评分学生设计海报,比较三维度表现
- AI在视觉表达上较准,但对共情与痛点识别一致性差
- 教师更信任人工评分,适合需要理解情境的创意评估
随着设计思维教育在中小学及高校的发展,评估融合视觉与文本的创意作品成为挑战。传统基于量表的评分依赖助教,在大规模多班级教学中耗时且不一致。本研究通过33位新加坡教育部教师参与的两轮活动,探索AI辅助评分与助教评分在设计思维教育中对学生海报评估的可靠性与准确性。比较了两个评分体系在共情与用户理解、痛点与机会识别、视觉传达三个维度的表现,并调查教师对AI评分、助教评分及混合评分的偏好。结果显示,教师与AI在共情与痛点识别上的统计一致性较低,视觉传达方面略有更高吻合度;在十组样本中,教师偏好助教评分的有六组。定性反馈指出AI在形成性反馈、评分一致性及促进学生自省方面具潜力,但难以捕捉上下文细节与创造性洞察。研究强调需构建融合计算效率与人类洞察的混合评估模式,推动创意领域负责任地应用AI,平衡自动化与人工判断以实现可扩展且符合教学目标的评价。
原文摘要 · Abstract (English)
As design thinking education grows in secondary and tertiary contexts, educators face the challenge of evaluating creative artefacts that combine visual and textual elements. Traditional rubric-based assessment is laborious, time-consuming, and inconsistent due to reliance on Teaching Assistants (TA) in large, multi-section cohorts. This paper presents an exploratory study investigating the reliability and perceived accuracy of AI-assisted assessment compared to TA-assisted assessment in evaluating student posters in design thinking education. Two activities were conducted with 33 Ministry of Education (MOE) Singapore school teachers to (1) compare AI-generated scores with TA grading across three key dimensions: empathy and user understanding, identification of pain points and opportunities, and visual communication, and (2) examine teacher preferences for AI-assigned, TA-assigned, and hybrid scores. Results showed low statistical agreement between instructor and AI scores for empathy and pain points, with slightly higher alignment for visual communication. Teachers preferred TA-assigned scores in six of ten samples. Qualitative feedback highlighted the potential of AI for formative feedback, consistency, and student self-reflection, but raised concerns about its limitations in capturing contextual nuance and creative insight. The study underscores the need for hybrid assessment models that integrate computational efficiency with human insights. This research contributes to the evolving conversation on responsible AI adoption in creative disciplines, emphasizing the balance between automation and human judgment for scalable and pedagogically sound assessment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。