PRISM提升生成式推荐的语义纯度与结构完整性,解决噪声干扰与信息丢失问题。
PRISM: Purified Representation and Integrated Semantic Modeling for Generative Sequential Recommendation
- 设计纯净语义量化器,通过自适应去噪和层级锚定增强词表鲁棒性。
- 引入动态语义融合机制,在稀疏场景下实现比基线显著更高的推荐准确率。
- 适合关注生成式推荐中语义建模与结构约束的研究者或工业应用开发者。
生成式序列推荐(GSR)将推荐任务重构为基于离散语义标识符(SIDs)的自回归序列生成,通常通过基于码本的量化获得。尽管该范式在统一检索与排序方面潜力巨大,现有框架仍存在两大局限:(1)语义标记不纯净且不稳定,量化方法难以应对交互噪声与码本坍塌,导致SIDs辨识模糊;(2)生成过程信息损失严重且结构薄弱,仅依赖粗粒度离散标记不可避免造成信息丢失,并忽略物品的层级逻辑。为此,本文提出新框架PRISM,包含纯净表示与集成语义建模。具体而言,为确保高质量标记化,设计了纯净语义量化器,通过自适应协同去噪与层级语义锚定构建鲁棒码本;为弥补量化中的信息损失,进一步提出集成语义推荐器,引入动态语义融合机制整合细粒度语义,并通过语义结构对齐目标强制逻辑有效性。PRISM在四个真实数据集上持续优于最先进基线,尤其在高稀疏场景下表现突出。
原文摘要 · Abstract (English)
Generative Sequential Recommendation (GSR) has emerged as a promising paradigm, reframing recommendation as an autoregressive sequence generation task over discrete Semantic IDs (SIDs), typically derived via codebook-based quantization. Despite its great potential in unifying retrieval and ranking, existing GSR frameworks still face two critical limitations: (1) impure and unstable semantic tokenization, where quantization methods struggle with interaction noise and codebook collapse, resulting in SIDs with ambiguous discrimination; and (2) lossy and weakly structured generation, where reliance solely on coarse-grained discrete tokens inevitably introduces information loss and neglects items' hierarchical logic. To address these issues, we propose a novel generative recommendation framework, PRISM, with Purified Representation and Integrated Semantic Modeling. Specifically, to ensure high-quality tokenization, we design a Purified Semantic Quantizer that constructs a robust codebook via adaptive collaborative denoising and hierarchical semantic anchoring mechanisms. To compensate for information loss during quantization, we further propose an Integrated Semantic Recommender, which incorporates a dynamic semantic integration mechanism to integrate fine-grained semantics and enforces logical validity through a semantic structure alignment objective. PRISM consistently outperforms state-of-the-art baselines across four real-world datasets, demonstrating substantial performance gains, particularly in high-sparsity scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。