arXiv:2604.19754cs.AIcs.LG2026-04

用合成数据提升Transformer模型对科学解释的评分准确率,解决类别不平衡问题。

Exploring Data Augmentation and Resampling Strategies for Transformer-Based Models to Address Class Imbalance in AI Scoring of Scientific Explanations in NGSS Classroom

论文配图:Exploring Data Augmentation and Resampling Strategies for Transformer-Based Models to Address Class Imbalance in AI Scoring of Scientific Explanations in NGSS Classroom
图 1 · 摘自论文原文
  • 采用GPT-4生成、ALP提取等策略增强数据,改善模型对稀有类别的识别
  • ALP方法在严重不平衡类别上实现精确率、召回率和F1值全为100%
  • 相比传统过采样,该方法更好保留初学者数据,适合教育测评场景

自动化评分可提供即时精准反馈,但基于NGSS学习进阶的科学解释评分中,高阶推理类别的类别不平衡仍是挑战。本研究基于1,466份高中生作答数据(11个二值化分析维度),使用SciBERT作为基线模型,探索三种数据增强策略:(1) GPT-4生成合成回答,(2) EASE词级提取过滤,(3) ALP句法级提取。微调虽提升召回率,但增强显著提升性能——GPT-4数据同时提高精确率与召回率,ALP在最严重不平衡的类别(5,6,7,9)上实现精确率、召回率与F1值均为100%。所有类别中EASE显著提升模型与人工评分的一致性,尤其在科学概念(1–6)与常见错误概念(7–11)上。与传统过采样(SMOTE)相比,该方法避免过拟合,保持学习进阶所需初级数据,为科学教育中的自动进阶评分提供可扩展解决方案。

原文摘要 · Abstract (English)

Automated scoring of students' scientific explanations offers the potential for immediate, accurate feedback, yet class imbalance in rubric categories particularly those capturing advanced reasoning remains a challenge. This study investigates augmentation strategies to improve transformer-based text classification of student responses to a physical science assessment based on an NGSS-aligned learning progression. The dataset consists of 1,466 high school responses scored on 11 binary-coded analytic categories. This rubric identifies six important components including scientific ideas needed for a complete explanation along with five common incomplete or inaccurate ideas. Using SciBERT as a baseline, we applied fine-tuning and test these augmentation strategies: (1) GPT-4--generated synthetic responses, (2) EASE, a word-level extraction and filtering approach, and (3) ALP (Augmentation using Lexicalized Probabilistic context-free grammar) phrase-level extraction. While fine-tuning SciBERT improved recall over baseline, augmentation substantially enhanced performance, with GPT data boosting both precision and recall, and ALP achieving perfect precision, recall, and F1 scores across most severe imbalanced categories (5,6,7 and 9). Across all rubric categories EASE augmentation substantially increased alignment with human scoring for both scientific ideas (Categories 1--6) and inaccurate ideas (Categories 7--11). We compared different augmentation strategies to a traditional oversampling method (SMOTE) in an effort to avoid overfitting and retain novice-level data critical for learning progression alignment. Findings demonstrate that targeted augmentation can address severe imbalance while preserving conceptual coverage, offering a scalable solution for automated learning progression-aligned scoring in science education.

AI评分数据增强教育测评文本分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。