arXiv:2511.20673cs.CLcs.IR2025-11被引 3

用双代码本区分热门与冷门物品,提升推荐精度与泛化能力

Semantics Meet Signals: Dual Codebook Representationl Learning for Generative Recommendation

  • 根据物品流行度动态分配协同过滤与语义代码本的令牌资源
  • 在多个数据集上优于基线模型,尤其改善长尾物品推荐效果
  • 适合需要兼顾热门与冷门物品推荐的工业级推荐系统

生成式推荐近年来成为统一检索与生成的强大范式,将物品表示为离散语义标记,并通过自回归模型实现灵活序列建模。尽管取得成功,现有方法依赖单一、统一的代码本编码所有物品,忽略了热门物品丰富的协同信号与长尾物品对语义理解的依赖性。我们提出FlexCode,一种流行度感知框架,自适应地在协同过滤(CF)代码本与语义代码本之间分配固定令牌预算。轻量级MoE动态平衡CF特异性精度与语义泛化,同时通过对齐与平滑目标维持跨流行度谱的一致性。在公开与工业规模数据集上的实验表明,FlexCode持续优于强基线。该方法为生成式推荐中的令牌表示提供新机制,在准确率与长尾鲁棒性上表现更优,为基于令牌的推荐模型中记忆与泛化的权衡提供了新视角。

原文摘要 · Abstract (English)

Generative recommendation has recently emerged as a powerful paradigm that unifies retrieval and generation, representing items as discrete semantic tokens and enabling flexible sequence modeling with autoregressive models. Despite its success, existing approaches rely on a single, uniform codebook to encode all items, overlooking the inherent imbalance between popular items rich in collaborative signals and long-tail items that depend on semantic understanding. We argue that this uniform treatment limits representational efficiency and hinders generalization. To address this, we introduce FlexCode, a popularity-aware framework that adaptively allocates a fixed token budget between a collaborative filtering (CF) codebook and a semantic codebook. A lightweight MoE dynamically balances CF-specific precision and semantic generalization, while an alignment and smoothness objective maintains coherence across the popularity spectrum. We perform experiments on both public and industrial-scale datasets, showing that FlexCode consistently outperform strong baselines. FlexCode provides a new mechanism for token representation in generative recommenders, achieving stronger accuracy and tail robustness, and offering a new perspective on balancing memorization and generalization in token-based recommendation models.

生成推荐双代码本长尾推荐自回归建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。