arXiv:2505.01442cs.IR2025-05

用性能空间框架科学选数据集,避免盲目跟风热门数据

Algorithm Performance Spaces for Strategic Dataset Selection

  • 构建算法性能空间,通过实测表现区分数据集差异
  • 提出三项指标量化数据集适用性,验证了不同规模电影评分数据的相似性
  • 适合想避免数据集偏见、追求严谨评估的研究者

推荐系统新算法的评估常依赖公开数据集(如MovieLens或Amazon)。但部分数据集因历史地位被过度使用,而非最适合当前研究场景。本文提出算法性能空间框架,基于算法在数据集上的表现差异来区分数据集。通过实验设计三项指标,用于量化并论证数据集选择的合理性。这些指标也验证了不同规模MovieLens数据集间的相似性假设。借助该框架与指标,实现了数据集的有效区分,并发现多样化的合理数据集组合。结果证明该框架有潜力,同时讨论了未来适配不同应用场景的研究方向。

原文摘要 · Abstract (English)

The evaluation of new algorithms in recommender systems frequently depends on publicly available datasets, such as those from MovieLens or Amazon. Some of these datasets are being disproportionately utilized primarily due to their historical popularity as baselines rather than their suitability for specific research contexts. This thesis addresses this issue by introducing the Algorithm Performance Space, a novel framework designed to differentiate datasets based on the measured performance of algorithms applied to them. An experimental study proposes three metrics to quantify and justify dataset selection to evaluate new algorithms. These metrics also validate assumptions about datasets, such as the similarity between MovieLens datasets of varying sizes. By creating an Algorithm Performance Space and using the proposed metrics, differentiating datasets was made possible, and diverse dataset selections could be found. While the results demonstrate the framework's potential, further research proposals and implications are discussed to develop Algorithm Performance Spaces tailored to diverse use cases.

推荐系统数据集评估算法评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。