arXiv:2504.10545cs.IRcs.LG2025-04中稿 · KDD

轻量对比文本嵌入提升生成推荐系统性能

HSTU-BLaIR: Lightweight Contrastive Text Embedding for Generative Recommender

  • 用轻量对比模型BLaIR增强序列建模的推荐框架
  • 在亚马逊和蒸汽平台数据集上表现优于大模型嵌入
  • 适合资源有限但需高效语义理解的推荐场景

近年来,生成建模与预训练语言模型在推荐系统中展现出互补优势。本文提出HSTU-BLaIR,一种将层次化序列转换单元(HSTU)生成推荐器与轻量级对比文本嵌入模型BLaIR相结合的混合框架。该框架通过文本元数据引入语义信号,同时保留HSTU强大的序列建模能力。我们在两个电商数据集上评估:亚马逊评论2023数据集的三个子集和Steam数据集。与原始HSTU推荐器及采用OpenAI最先进的text-embedding-3-large模型嵌入的变体相比,尽管后者参数更多、训练语料更广,但本方法在几乎所有指标上表现更优。具体而言,HSTU-BLaIR在除一项指标外的所有指标上均超越该变体,且在另一项指标上持平。结果表明,对比文本嵌入在计算高效的推荐场景中具有显著有效性。

原文摘要 · Abstract (English)

Recent advances in recommender systems have underscored the complementary strengths of generative modeling and pretrained language models. We propose HSTU-BLaIR, a hybrid framework that augments the Hierarchical Sequential Transduction Unit (HSTU)-based generative recommender with BLaIR, a lightweight contrastive text embedding model. This integration enriches item representations with semantic signals from textual metadata while preserving HSTU's powerful sequence modeling capabilities. We evaluate HSTU-BLaIR on two e-commerce datasets: three subsets from the Amazon Reviews 2023 dataset and the Steam dataset. We compare its performance against both the original HSTU-based recommender and a variant augmented with embeddings from OpenAI's state-of-the-art \texttt{text-embedding-3-large} model. Despite the latter being trained on a substantially larger corpus with significantly more parameters, our lightweight BLaIR-enhanced approach -- pretrained on domain-specific data -- achieves better performance in nearly all cases. Specifically, HSTU-BLaIR outperforms the OpenAI embedding-based variant on all but one metric, where it is marginally lower, and matches it on another. These findings highlight the effectiveness of contrastive text embeddings in compute-efficient recommendation settings.

生成推荐对比学习轻量模型文本嵌入

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。