构建首个推荐系统全能评估基准,推动推荐模型向通用智能演进
OpenOneRec Technical Report
- 设计涵盖8类任务的RecIF-Bench,覆盖从预测到推理的完整能力链
- 释放9600万用户交互数据集与可复现训练框架,支持大规模实验
- 推出1.7B/8B参数的OneRec-Foundation模型,在多场景下召回率提升26.8%
尽管OneRec系列已将碎片化的推荐流程统一为端到端生成框架,但推荐系统与通用智能之间仍存在显著差距。受限于孤立数据,现有系统仅能作为领域专家——擅长模式匹配,却缺乏世界知识、推理能力与指令遵循能力。这一局限进一步因缺乏综合评估基准而加剧。为此,本工作贡献如下:1)提出RecIF-Bench,一个涵盖8种多样化任务的综合性评估基准,全面检验从基础预测到复杂推理的能力;同时发布包含16万用户9600万条交互记录的大规模训练数据集,支持可复现研究。2)构建完整开源训练框架,涵盖数据处理、联合预训练与后训练流程;基于该框架,证明推荐能力可实现可预测扩展,并有效缓解通用知识灾难性遗忘。3)发布OneRec Foundation(1.7B与8B参数),在RecIF-Bench所有任务上刷新当前最优性能;迁移至Amazon基准时,其平均Recall@10在10个不同数据集上相较最强基线提升26.8%。此项工作标志着迈向真正智能推荐系统的重要一步。然而,实现该愿景仍面临重大技术与理论挑战,亟需更广泛的研究投入。
原文摘要 · Abstract (English)
While the OneRec series has successfully unified the fragmented recommendation pipeline into an end-to-end generative framework, a significant gap remains between recommendation systems and general intelligence. Constrained by isolated data, they operate as domain specialists-proficient in pattern matching but lacking world knowledge, reasoning capabilities, and instruction following. This limitation is further compounded by the lack of a holistic benchmark to evaluate such integrated capabilities. To address this, our contributions are: 1) RecIF Bench & Open Data: We propose RecIF-Bench, a holistic benchmark covering 8 diverse tasks that thoroughly evaluate capabilities from fundamental prediction to complex reasoning. Concurrently, we release a massive training dataset comprising 96 million interactions from 160,000 users to facilitate reproducible research. 2) Framework & Scaling: To ensure full reproducibility, we open-source our comprehensive training pipeline, encompassing data processing, co-pretraining, and post-training. Leveraging this framework, we demonstrate that recommendation capabilities can scale predictably while mitigating catastrophic forgetting of general knowledge. 3) OneRec-Foundation: We release OneRec Foundation (1.7B and 8B), a family of models establishing new state-of-the-art (SOTA) results across all tasks in RecIF-Bench. Furthermore, when transferred to the Amazon benchmark, our models surpass the strongest baselines with an average 26.8% improvement in Recall@10 across 10 diverse datasets (Figure 1). This work marks a step towards building truly intelligent recommender systems. Nonetheless, realizing this vision presents significant technical and theoretical challenges, highlighting the need for broader research engagement in this promising direction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。