arXiv:2412.10401cs.LG2024-12NeurIPS被引 5

构建大规模儿童阅读数据集,用自监督学习预测早期阅读风险

Scalable Early Childhood Reading Performance Prediction

  • 用掩码输入预训练MLP,实现自监督学习
  • 在44所学校6916名学生数据上准确识别早期阅读模式
  • 适合教育科技、个性化教学研究者使用

学生阅读表现模型可帮助教育者提前识别高危学生,实现早期精准干预。然而,目前缺乏适合建模和预测未来阅读表现的公开教育数据集。本文提出增强型核心阅读教学(ECRI)数据集,为跨44所学校、6916名学生和172名教师的大规模纵向表格数据集。我们利用该数据集实证评估了前沿机器学习模型在多变量与不完整测量中识别早期教育模式的能力。具体而言,我们展示了一种简单自监督策略:通过掩码输入对多层感知机(MLP)网络进行预训练,其性能优于多个强基线模型,并在多样教育环境中具备良好泛化能力。为促进精准建模及负责任的个体化早期干预策略发展,我们的数据与代码已公开于 https://ecri-data.github.io/。

原文摘要 · Abstract (English)

Models for student reading performance can empower educators and institutions to proactively identify at-risk students, thereby enabling early and tailored instructional interventions. However, there are no suitable publicly available educational datasets for modeling and predicting future reading performance. In this work, we introduce the Enhanced Core Reading Instruction ECRI dataset, a novel large-scale longitudinal tabular dataset collected across 44 schools with 6,916 students and 172 teachers. We leverage the dataset to empirically evaluate the ability of state-of-the-art machine learning models to recognize early childhood educational patterns in multivariate and partial measurements. Specifically, we demonstrate a simple self-supervised strategy in which a Multi-Layer Perception (MLP) network is pre-trained over masked inputs to outperform several strong baselines while generalizing over diverse educational settings. To facilitate future developments in precise modeling and responsible use of models for individualized and early intervention strategies, our data and code are available at https://ecri-data.github.io/.

阅读预测教育数据自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。