arXiv:2412.02108cs.LGcs.CY2024-12被引 11

对比21种数据增强方法,发现SMOTE-ENN最有效且提速近半。

Evaluating the Impact of Data Augmentation on Predictive Model Performance

  • 在学术表现预测任务中系统测试21种数据增强技术。
  • SMOTE-ENN使平均AUC提升0.01,训练时间减半。
  • 部分方法反而降低性能,建议谨慎使用。

在监督学习研究中,大规模训练数据对结果有效性至关重要。然而,学习分析(LA)领域获取原始数据颇具挑战。数据增强可通过扩展和多样化数据缓解此问题,但其在LA中的应用仍不充分。本文系统比较了多种数据增强技术对典型LA任务——学术成果预测性能的影响。实验基于前人LAK研究复现了四种SML模型,并以AUC值为评估标准。在21种增强技术中,SMOTE-ENN采样表现最优,平均AUC提升0.01,训练时间约为基线的一半。此外,我们还测试了99种技术组合,发现向SMOTE-ENN添加噪声可带来约0.014的微小但显著的性能提升。值得注意的是,部分增强方法显著降低了预测性能或加剧了由随机性引发的波动。本研究贡献有二:首先,实证表明采样类方法在LA中提供最可靠、计算更高效的性能改进,优于复杂超参设置的生成式方法;其次,强调通过独立复现验证近期研究的重要性。

原文摘要 · Abstract (English)

In supervised machine learning (SML) research, large training datasets are essential for valid results. However, obtaining primary data in learning analytics (LA) is challenging. Data augmentation can address this by expanding and diversifying data, though its use in LA remains underexplored. This paper systematically compares data augmentation techniques and their impact on prediction performance in a typical LA task: prediction of academic outcomes. Augmentation is demonstrated on four SML models, which we successfully replicated from a previous LAK study based on AUC values. Among 21 augmentation techniques, SMOTE-ENN sampling performed the best, improving the average AUC by 0.01 and approximately halving the training time compared to the baseline models. In addition, we compared 99 combinations of chaining 21 techniques, and found minor, although statistically significant, improvements across models when adding noise to SMOTE-ENN (+0.014). Notably, some augmentation techniques significantly lowered predictive performance or increased performance fluctuation related to random chance. This paper's contribution is twofold. Primarily, our empirical findings show that sampling techniques provide the most statistically reliable performance improvements for LA applications of SML, and are computationally more efficient than deep generation methods with complex hyperparameter settings. Second, the LA community may benefit from validating a recent study through independent replication.

数据增强学习分析模型性能SMOTE-ENN

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。