统一生成与排序的语义ID框架,提升电商推荐精准度
UniSGR: Unified Framework for Semantic ID Generation and Ranking
- 两阶段训练:跨场景预训练+场景化对齐优化
- 生成与排序联合优化,显著提升多目标推荐效果
- 新推理策略突破传统束搜索效率瓶颈,适合大规模推荐
推荐系统在现代电商平台中至关重要。尽管生成式召回已成为缓解多阶段级联架构局限性的有前景范式,现有方法仍难以应对细粒度多目标排序挑战。为此,我们提出UniSGR——一种统一的语义ID生成与排序框架。UniSGR采用两阶段训练:第一阶段为多场景预训练,利用混合业务场景数据学习;第二阶段为场景特定对齐,联合优化价值感知并行多标记预测(VA-PMTP)与统一多目标排序模块。为更好对齐生成与下游排序,引入由漏斗感知对比学习引导的任务感知标记(TAT)。此外,提出语义树注意力与重组键值缓存(STARK)推理策略,消除传统束搜索中的关键效率瓶颈。在大型电商平台上的大量离线实验验证了UniSGR的有效性与可扩展性。
原文摘要 · Abstract (English)
Recommendation systems play a pivotal role in modern e-commerce platforms. While generative retrieval has emerged as a promising paradigm for alleviating the limitations of multi-stage cascade architectures, existing methods still struggle with fine-grained multi-objective ranking. To bridge this gap, we propose UniSGR, a Unified framework for Semantic ID Generation and Ranking. UniSGR adopts a two-stage training paradigm: a multi-scenario pre-training stage that learns from mixed business-scenario data, followed by a scenario-specific alignment stage that jointly optimizes Value-Aware Parallel Multi-Token Prediction (VA-PMTP) and a unified multi-objective ranking module. To better align generation with downstream ranking, we introduce Task-Aware Tokens (TAT) guided by Funnel-Aware Contrastive Learning. Furthermore, we propose Semantic Tree Attention with Reorganized KV cache (STARK), an inference strategy that removes key efficiency bottlenecks in conventional beam search. Extensive offline experiments on a large-scale e-commerce platform demonstrate the effectiveness and scalability of UniSGR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。