arXiv:2409.05526cs.IR2024-09被引 2

RBoard统一评测推荐系统,让实验可复现、结果可比。

RBoard: A Unified Platform for Reproducible and Reusable Recommender System Benchmarks

  • 构建统一平台,覆盖点击率预测、顶N推荐等任务
  • 多数据集评测+标准化流程,确保结果可比性
  • 开源代码可下载,适合学术与工业界研究复现

推荐系统研究缺乏可复现和算法对比的标准基准。我们提出RBoard,一个全新框架,为多种推荐任务(包括点击率预测、顶N推荐等)提供全面的基准测试平台。其核心目标是实现跨场景的完全可复现与可重用实验。框架在每个任务下对多种数据集评估算法,并聚合结果以进行整体性能分析。通过实施标准化评估协议,确保一致性与可比性。为提升可复现性,所有用户提交的代码均可轻松下载并执行,使研究者能可靠复现成果并在此基础上推进工作。通过提供统一平台,支持各类推荐场景下的严谨、可复现评估,RBoard旨在加速领域进展,并在学术界与工业界建立新的推荐系统基准标准。平台地址:https://rboard.org,演示视频:https://bit.ly/rboard-demo。

原文摘要 · Abstract (English)

Recommender systems research lacks standardized benchmarks for reproducibility and algorithm comparisons. We introduce RBoard, a novel framework addressing these challenges by providing a comprehensive platform for benchmarking diverse recommendation tasks, including CTR prediction, Top-N recommendation, and others. RBoard's primary objective is to enable fully reproducible and reusable experiments across these scenarios. The framework evaluates algorithms across multiple datasets within each task, aggregating results for a holistic performance assessment. It implements standardized evaluation protocols, ensuring consistency and comparability. To facilitate reproducibility, all user-provided code can be easily downloaded and executed, allowing researchers to reliably replicate studies and build upon previous work. By offering a unified platform for rigorous, reproducible evaluation across various recommendation scenarios, RBoard aims to accelerate progress in the field and establish a new standard for recommender systems benchmarking in both academia and industry. The platform is available at https://rboard.org and the demo video can be found at https://bit.ly/rboard-demo.

推荐系统可复现性基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。