arXiv:2410.22972cs.IR2024-10被引 9

DataRec统一推荐系统数据管理,提升实验可复现性。

DataRec: A Python Library for Standardized and Reproducible Data Management in Recommender Systems

  • 提供标准化数据处理流程,支持版本控制与框架集成
  • 基于55篇顶会论文设计,覆盖主流推荐场景
  • 适合希望提升实验透明度的研究者与工业开发者

推荐系统在多个领域展现出显著影响,但实验结果的可复现性仍是长期挑战。主要障碍在于预处理阶段的数据管理碎片化且不透明,数据集选择、过滤与划分策略对结果影响显著。为此,我们提出DataRec——一个开源的Python库,专为统一和简化推荐系统研究中的数据处理而设计。通过提供可复现的数据准备、数据版本管理及与其他框架的无缝集成,DataRec推动方法学标准化、互操作性与实验可比性。其设计基于对55篇顶尖推荐系统研究的深入分析,采纳最佳实践并规避常见数据管理陷阱。最终,该工作促进公平基准测试,增强实验可信度,提升推荐系统社区的整体研究质量。DataRec库、文档与示例已公开:https://github.com/sisinflab/DataRec。

原文摘要 · Abstract (English)

Recommender systems have demonstrated significant impact across diverse domains, yet ensuring the reproducibility of experimental findings remains a persistent challenge. A primary obstacle lies in the fragmented and often opaque data management strategies employed during the preprocessing stage, where decisions about dataset selection, filtering, and splitting can substantially influence outcomes. To address these limitations, we introduce DataRec, an open-source Python-based library specifically designed to unify and streamline data handling in recommender system research. By providing reproducible routines for dataset preparation, data versioning, and seamless integration with other frameworks, DataRec promotes methodological standardization, interoperability, and comparability across different experimental setups. Our design is informed by an in-depth review of 55 state-of-the-art recommendation studies ensuring that DataRec adopts best practices while addressing common pitfalls in data management. Ultimately, our contribution facilitates fair benchmarking, enhances reproducibility, and fosters greater trust in experimental results within the broader recommender systems community. The DataRec library, documentation, and examples are freely available at https://github.com/sisinflab/DataRec.

推荐系统数据管理可复现性Python库

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。