arXiv:2512.10388cs.IRcs.AI2025-12被引 3

融合语义与哈希ID,提升长尾推荐效果

The Best of the Two Worlds: Harmonizing Semantic and Hash IDs for Sequential Recommendation

  • 双分支架构并行建模语义ID与哈希ID
  • 在三个公开数据集和真实平台均超越基线
  • 特别适合长尾物品多的推荐场景

传统序列推荐系统通常使用唯一哈希ID(HID)构建物品嵌入,主要捕捉历史用户-物品交互中的协同信号,但在长尾场景下表现不佳。引入辅助信息的方法常受共现信号噪声或密集嵌入语义同质性影响。相比之下,语义ID(SID)支持代码共享和多粒度语义建模,但现有方法存在协同信号过强问题:常用量化机制削弱了头段物品标识符的独特性,导致头尾物品性能权衡。为此,我们提出H2Rec框架,通过双分支建模同时保留SID的多粒度语义与HID的唯一协同身份,并设计双层对齐策略实现两者的有效知识迁移与鲁棒偏好建模。在三个公开基准和一个大规模商用平台上的离线与在线实验表明,H2Rec在头尾物品推荐质量间取得更好平衡,持续优于现有基线。

原文摘要 · Abstract (English)

Conventional Sequential Recommender Systems (SRS) typically assign unique hash IDs (HID) to construct item embeddings, which mainly capture collaborative signals from historical user-item interactions. However, such embeddings are vulnerable in long-tail scenarios where most items are rarely consumed. Recent methods that incorporate auxiliary information often face noisy collaborative sharing from co-occurrence signals or semantic homogeneity caused by flat dense embeddings. In contrast, Semantic IDs (SID), with their support for code sharing and multi-granular semantic modeling, offer a promising alternative. Nevertheless, SID-based methods are hindered by a collaborative overwhelming phenomenon: commonly adopted quantization mechanisms compromise the identifier uniqueness needed to model head items, resulting in a performance trade-off between head and tail items. To address this challenge, we propose H2Rec, a novel framework that harmonizes SID and HID. We design a dual-branch modeling architecture that simultaneously captures the multi-granular semantics of SID while preserving the unique collaborative identity provided by HID. Moreover, we introduce a dual-level alignment strategy to bridge the two representations, enabling effective knowledge transfer and robust preference modeling. Extensive offline experiments on three public benchmarks and online experiments on a large-scale commercial platform demonstrate that H2Rec achieves a better balance between head and tail recommendation quality and consistently outperforms existing baselines.

序列推荐长尾问题语义编码双分支模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。