arXiv:2602.17314cs.CYcs.DB2026-02被引 1

梳理学习分析领域开源数据现状,提出8项可操作的数据发布指南。

Open Datasets in Learning Analytics: Trends, Challenges, and Best PRACTICE

  • 系统调研1125篇顶会论文,发现204篇使用了172个开源数据集。
  • 超80%数据集未被以往调查覆盖,揭示领域数据开放严重不足。
  • 提出PRACTICE指南,助研究者高效、合规地发布数据。

开源数据在学习分析、教育数据挖掘和人工智能教育三大交叉领域至关重要。本研究对过去五年内学习分析领域三大旗舰会议(LAK、EDM、AIED)的1,125篇论文进行人工核查,发现并分析了204篇论文中使用的172个开源数据集。其中143个数据集未被此前任何调查收录。研究系统梳理了数据集的应用场景、分析方法与特性,揭示当前数据开放中的关键缺口。基于调研结果,提出以PRACTICE为缩写的八项实践建议与检查清单,指导研究者规范发布数据。同时公开原始数据集:一个包含所有发现数据集及其对应论文的标注清单。研究旨在推动学习分析等领域的数据开放文化发展。

原文摘要 · Abstract (English)

Open datasets play a crucial role in three research domains that intersect data science and education: learning analytics, educational data mining, and artificial intelligence in education. Researchers in these domains apply computational methods to analyze data from educational contexts, aiming to better understand and improve teaching and learning. Providing open datasets alongside research papers supports reproducibility, collaboration, and trust in research findings. It also provides individual benefits for authors, such as greater visibility, credibility, and citation potential. Despite these advantages, the availability of open datasets and the associated practices within the learning analytics research communities, especially at their flagship conference venues, remain unclear. We surveyed available datasets published alongside research papers in learning analytics. We manually examined 1,125 papers from three flagship conferences (LAK, EDM, and AIED) over the past five years. We discovered, categorized, and analyzed 172 datasets used in 204 publications. Our study presents the most comprehensive collection and analysis of open educational datasets to date, along with the most detailed categorization. Of the 172 datasets identified, 143 were not captured in any prior survey of open data in learning analytics. We provide insights into the datasets' context, analytical methods, use, and other properties. Based on this survey, we summarize the current gaps in the field. Furthermore, we list practical recommendations, advice, and 8-item guidelines under the acronym PRACTICE with a checklist to help researchers publish their data. Lastly, we share our original dataset: an annotated inventory detailing the discovered datasets and the corresponding publications. We hope these findings will support further adoption of open data practices in learning analytics communities and beyond.

学习分析开源数据PRACTICE数据共享

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。