用学习进展自动生成科学作业反馈,效果不输专家人工设计。
Using Learning Progressions to Guide AI Feedback for Science Learning
- 基于学习进展自动生成评分标准,替代人工编写
- 207名中学生化学写作反馈质量无显著差异
- 适合教育AI规模化落地,尤其缺乏专家资源的场景
生成式人工智能可为形成性反馈提供可扩展支持,但现有AI反馈多依赖领域专家设计的任务专属评分标准。虽然有效,但评分标准编写耗时,限制了在不同教学情境中的推广。学习进展(LP)提供了学生理解能力发展过程的理论基础,或可作为替代方案。本研究考察基于学习进展的评分标准生成流程是否能产生与专家制定评分标准相当的AI反馈质量。分析了207名中学生在化学任务中撰写的科学解释所对应的AI反馈。比较两种流程:(a) 由人类专家设计的任务特定评分标准;(b) 在评分与反馈生成前,从学习进展自动生成的任务特定评分标准。两名人工评估者使用包含10个子维度的多维评分标准评估反馈质量,涵盖清晰度、准确性、相关性、参与度与动机、反思性。评分者间一致性高,百分比一致率介于89%至100%,科恩κ值为0.66至0.88。配对t检验显示,两种流程在清晰度(t1=0.00, p1=1.000;t2=0.84, p2=0.399)、相关性(t1=0.28, p1=0.782;t2=-0.58, p2=0.565)、参与度与动机(t1=0.50, p1=0.618;t2=-0.58, p2=0.565)及反思性(t=-0.45, p=0.656)方面均无统计学显著差异。结果表明,基于学习进展的评分标准生成流程可作为有效替代方案。
原文摘要 · Abstract (English)
Generative artificial intelligence (AI) offers scalable support for formative feedback, yet most AI-generated feedback relies on task-specific rubrics authored by domain experts. While effective, rubric authoring is time-consuming and limits scalability across instructional contexts. Learning progressions (LP) provide a theoretically grounded representation of students' developing understanding and may offer an alternative solution. This study examines whether an LP-driven rubric generation pipeline can produce AI-generated feedback comparable in quality to feedback guided by expert-authored task rubrics. We analyzed AI-generated feedback for written scientific explanations produced by 207 middle school students in a chemistry task. Two pipelines were compared: (a) feedback guided by a human expert-designed, task-specific rubric, and (b) feedback guided by a task-specific rubric automatically derived from a learning progression prior to grading and feedback generation. Two human coders evaluated feedback quality using a multi-dimensional rubric assessing Clarity, Accuracy, Relevance, Engagement and Motivation, and Reflectiveness (10 sub-dimensions). Inter-rater reliability was high, with percent agreement ranging from 89% to 100% and Cohen's kappa values for estimable dimensions (kappa = .66 to .88). Paired t-tests revealed no statistically significant differences between the two pipelines for Clarity (t1 = 0.00, p1 = 1.000; t2 = 0.84, p2 = .399), Relevance (t1 = 0.28, p1 = .782; t2 = -0.58, p2 = .565), Engagement and Motivation (t1 = 0.50, p1 = .618; t2 = -0.58, p2 = .565), or Reflectiveness (t = -0.45, p = .656). These findings suggest that the LP-driven rubric pipeline can serve as an alternative solution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。