arXiv:2603.00632cs.IRcs.LG2026-03KDD被引 11

区分碰撞类型,让推荐系统的短编码更精准

Stop Treating Collisions Equally: Qualification-Aware Semantic ID Learning for Recommendation at Industrial Scale

论文配图:Stop Treating Collisions Equally: Qualification-Aware Semantic ID Learning for Recommendation at Industrial Scale
图 1 · 摘自论文原文
  • 根据碰撞严重程度动态调整惩罚力度,避免误伤
  • 在公开数据集上提升排序效果5.9%,冷启动推荐订单量增6.42%
  • 适合大规模工业级推荐系统,可直接接入现有框架

语义ID(SIDs)是从多模态商品特征中提取的紧凑离散表示,作为基于ID与生成式推荐的统一抽象。但高质量SIDs的学习面临两大挑战:(1) 碰撞问题——量化后的令牌空间易发生语义不同商品被分配相同或高度相似的SID组合,导致语义混淆;(2) 碰撞信号异质性——并非所有碰撞都等效有害,部分反映语义无关商品的真实冲突,另一些则源于良性冗余或系统性数据效应。为此,我们提出资格感知语义ID学习(QuaSID),一种端到端框架,通过选择性排斥有资格的冲突对,并按碰撞严重性调节排斥强度来学习带资格标签的SIDs。QuaSID包含两个机制:汉明引导的边际排斥,将低汉明距离的SID重叠转化为编码空间中显式的、严重性加权的几何约束;以及冲突感知有效对掩码,屏蔽协议引起的良性重叠以净化排斥监督信号。此外,引入双塔对比目标,将协同信号注入分词过程。在公开基准和工业数据上的实验验证了QuaSID的有效性。在公开数据集上,其始终优于强基线,使顶级K排名质量提升5.9%,同时增加SID组成多样性。在快手电商的线上A/B测试中(5%流量),其使排序GMV-S2提升2.38%,冷启动检索完成订单量最高提升6.42%。最后,我们证明该排斥损失具有即插即用性,可在多个数据集上增强多种SID学习框架。

原文摘要 · Abstract (English)

Semantic IDs (SIDs) are compact discrete representations derived from multimodal item features, serving as a unified abstraction for ID-based and generative recommendation. However, learning high-quality SIDs remains challenging due to two issues. (1) Collision problem: the quantized token space is prone to collisions, in which semantically distinct items are assigned identical or overly similar SID compositions, resulting in semantic entanglement. (2) Collision-signal heterogeneity: collisions are not uniformly harmful. Some reflect genuine conflicts between semantically unrelated items, while others stem from benign redundancy or systematic data effects. To address these challenges, we propose Qualification-Aware Semantic ID Learning (QuaSID), an end-to-end framework that learns collision-qualified SIDs by selectively repelling qualified conflict pairs and scaling the repulsion strength by collision severity. QuaSID consists of two mechanisms: Hamming-guided Margin Repulsion, which translates low-Hamming SID overlaps into explicit, severity-scaled geometric constraints on the encoder space; and Conflict-Aware Valid Pair Masking, which masks protocol-induced benign overlaps to denoise repulsion supervision. In addition, QuaSID incorporates a dual-tower contrastive objective to inject collaborative signals into tokenization. Experiments on public benchmarks and industrial data validate QuaSID. On public datasets, QuaSID consistently outperforms strong baselines, improving top-K ranking quality by 5.9% over the best baseline while increasing SID composition diversity. In an online A/B test on Kuaishou e-commerce with a 5% traffic split, QuaSID increases ranking GMV-S2 by 2.38% and improves completed orders on cold-start retrieval by up to 6.42%. Finally, we show that the proposed repulsion loss is plug-and-play and enhances a range of SID learning frameworks across datasets.

推荐系统语义编码工业级应用对抗碰撞

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。