arXiv:2509.03131cs.IRcs.LG2025-09EMNLP被引 5

用统一物品编码预训练大模型,实现跨域推荐零样本效果突破

RecBase: Generative Foundation Model Pretraining for Zero-Shot Recommendation

  • 构建面向推荐任务的通用预训练模型,用物品层级编码提升跨域泛化
  • 1.5B参数模型在8个数据集上超越7B参数大模型的零样本推荐表现
  • 适合需要跨域推荐且无标注数据的场景,如新平台冷启动

基于大语言模型的推荐系统虽有进展,但因语言预训练与推荐任务本质不匹配,跨域泛化能力受限。现有方法依赖语言级知识,难以捕捉跨域动态的物品级用户兴趣。为此,我们提出RecBase,一种以推荐为导向、无领域偏见的基础模型。RecBase利用大规模异构跨域语料库,通过统一文本表示和特征映射增强跨域泛化能力。为对齐不同域的物品语义,引入统一物品分词器,将物品编码为层次化概念标识符,实现结构化表示与高效词表共享。模型采用自回归目标训练,以捕捉复杂的物品级序列模式。在8个真实世界数据集上,我们的1.5B参数模型在零样本与跨域推荐任务中达到或超过7B参数大模型的性能。

原文摘要 · Abstract (English)

Recent advances in LLM-based recommendation have shown promise, yet their cross-domain generalization is hindered by a fundamental mismatch between language-centric pretraining and the recommendation task. Existing methods, relying on language-level knowledge, fail to capture dynamic, item-level user interests across domains. To bridge this gap, we propose RecBase, a domain-agnostic foundational model pretrained with a recommendation-oriented objective. RecBase leverages a large-scale, heterogeneous, cross-domain corpus with unified textual representations and feature mappings to enhance cross-domain generalization. To further align item semantics across domains, we introduce a unified item tokenizer that encodes items into hierarchical concept identifiers, enabling structured representation and efficient vocabulary sharing. The model is trained using an autoregressive objective to capture complex item-level sequential patterns. On eight real-world datasets, our 1.5B-parameter model matches or surpasses the performance of LLM baselines up to 7B parameters in zero-shot and cross-domain recommendation tasks.

推荐系统基础模型零样本跨域推荐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。