arXiv:2510.26095cs.IRcs.CL2025-10NeurIPS被引 1

构建可复现的推荐系统评测基准,用真实网页浏览数据检验模型泛化能力。

ORBIT -- Open Recommendation Benchmark for Reproducible Research with Hidden Tests

  • 统一公开数据集与可复现划分,提供透明评估框架。
  • 引入8700万网页的网页推荐任务ClueWeb-Reco,作为隐藏测试集。
  • 验证现有模型在大规模推荐中不足,提示LLM融合潜力。

推荐系统是影响深远的AI应用,每日服务数十亿用户,为其提供个性化内容推荐。然而,当前研究受限于无法反映真实用户行为的数据集和不一致的评估设置,导致结论模糊。本文提出开放推荐基准ORBIT(Open Recommendation Benchmark for Reproducible Research with Hidden Tests),提供标准化评估框架,包含公开数据集、可复现的训练/测试划分及透明的公共排行榜。新增网页推荐任务ClueWeb-Reco,基于8700万条经用户同意、隐私保障的真实网页浏览序列构建,模拟现代推荐场景。该数据集保留为排行榜的隐藏测试集,用于检验模型泛化能力。我们在公开数据集上评估12种代表性推荐模型,并在隐藏测试集引入提示式LLM基线。公开结果反映推荐系统整体进步,但个体表现差异大;隐藏测试结果揭示现有方法在大规模网页推荐中的局限性,凸显结合LLM的改进潜力。ORBIT基准、排行榜与代码库已开源:https://www.open-reco-bench.ai。

原文摘要 · Abstract (English)

Recommender systems are among the most impactful AI applications, interacting with billions of users every day, guiding them to relevant products, services, or information tailored to their preferences. However, the research and development of recommender systems are hindered by existing datasets that fail to capture realistic user behaviors and inconsistent evaluation settings that lead to ambiguous conclusions. This paper introduces the Open Recommendation Benchmark for Reproducible Research with HIdden Tests (ORBIT), a unified benchmark for consistent and realistic evaluation of recommendation models. ORBIT offers a standardized evaluation framework of public datasets with reproducible splits and transparent settings for its public leaderboard. Additionally, ORBIT introduces a new webpage recommendation task, ClueWeb-Reco, featuring web browsing sequences from 87 million public, high-quality webpages. ClueWeb-Reco is a synthetic dataset derived from real, user-consented, and privacy-guaranteed browsing data. It aligns with modern recommendation scenarios and is reserved as the hidden test part of our leaderboard to challenge recommendation models' generalization ability. ORBIT measures 12 representative recommendation models on its public benchmark and introduces a prompted LLM baseline on the ClueWeb-Reco hidden test. Our benchmark results reflect general improvements of recommender systems on the public datasets, with variable individual performances. The results on the hidden test reveal the limitations of existing approaches in large-scale webpage recommendation and highlight the potential for improvements with LLM integrations. ORBIT benchmark, leaderboard, and codebase are available at https://www.open-reco-bench.ai.

推荐系统可复现性基准测试LLM融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。