通过元数据增强提升软件挖掘数据集的可发现性与可重用性
MIRAGE: Metadata-Integrated Repository Analysis and Guided Enhancement for MSR Datasets

- 基于语义学者API收集2013-2024年论文元数据,扩展数据集目录
- 发现托管平台和数据格式影响引用模式与可用性
- 支持更高效的数据集复用与研究资产评估
本文提出一种改进的软件挖掘数据集(MSR)分析方法,通过元数据增强、FAIR性评估与主题驱动分析实现。研究在早期专用于MSR数据集分析的目录基础上,新增数据集注释,丰富元数据类别,并提供更高级的筛选功能。利用语义学者API收集2013至2024年相关论文的元数据,基于隐含狄利克雷分布(LDA)主题建模与统计分析进行研究。扩充后的数据集目录包含仓库托管平台、格式、可访问性、可重用性及数据质量等属性。研究发现,托管平台和数据格式显著影响引用模式与数据集可用性。增强的标注方法提升了对MSR数据集的分析能力与可发现性,有助于研究成果的更有效复用与评估。
原文摘要 · Abstract (English)
This paper proposes an improved approach to the analysis of Mining Software Repositories (MSR) datasets via metadata enrichment, FAIRness assessment, and topic-driven analysis. This research expands upon an earlier dataset directory created specifically for the analysis of MSR datasets by adding new annotations to the datasets, enriching the metadata categories, and offering more advanced filtering options. The metadata of the MSR papers presented from 2013 to 2024 has been gathered using the Semantic Scholar API. The analysis is based on Latent Dirichlet Allocation (LDA) topic modeling and statistical analysis. Dataset-level attributes were included into the expanded dataset directory, namely repository hosting site, format, accessibility, reusability, and dataset quality. The study reveals that the choice of repository hosting sites and data formats influences citation patterns and dataset usability. Furthermore, the enhanced annotation approach improves the analysis and discoverability of MSR datasets, supporting more effective reuse and evaluation of research artifacts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。