arXiv:2604.23522cs.IRcs.MM2026-04被引 4

提出自适应语义ID学习框架,解决推荐系统中编码冲突问题

Beyond Static Collision Handling: Adaptive Semantic ID Learning for Multimodal Recommendation at Industrial Scale

论文配图:Beyond Static Collision Handling: Adaptive Semantic ID Learning for Multimodal Recommendation at Industrial Scale
图 1 · 摘自论文原文
  • 通过两阶段调节机制,智能区分应抑制或保留的编码重叠
  • 在公开数据集上召回率和NDCG提升约4.5%,工业测试中GMV增0.98%
  • 适合大规模多模态推荐场景,尤其关注编码紧凑性与语义保真

现代推荐系统涉及海量多模态物品,可扩展的物品标识需平衡紧凑性、语义保真度与下游效果。语义ID(SIDs)通过从多模态信号生成短离散令牌序列来表示物品,为检索、排序和生成推荐提供紧凑接口。但有效SID学习受限于碰撞问题,即不同物品被分配相同或高度混淆的编码。现有方法主要依赖改进量化或固定重叠正则化,无法自适应判断是否应抑制或保留重叠。本文提出AdaSID,一种用于推荐的自适应语义ID学习框架。AdaSID通过两阶段过程调节SID重叠:首先,在涉及物品语义兼容时放松排斥,保留可接受的共享而非统一分离所有碰撞;其次,根据局部碰撞密度和训练进度分配剩余调节压力,在高密度区域加强控制,同时逐步将优化重心转向推荐对齐。该设计自适应决定哪些重叠应惩罚、惩罚强度及何时转移学习重点。大量离线与在线实验验证了其有效性。在两个公开基准上,AdaSID平均相比强基线提升约4.5%的召回率与NDCG,同时提升代码本利用率与SID多样性。在快手电商场景中,针对数千万用户的短视频检索在线A/B测试取得显著收益,包括0.98%的GMV提升,工业排名评估也显示一致的AUC改善。

原文摘要 · Abstract (English)

Modern recommendation systems involve massive catalogs of multimodal items, where scalable item identification must balance compactness, semantic fidelity, and downstream effectiveness. Semantic IDs (SIDs) address this need by representing items as short discrete token sequences derived from multimodal signals, providing a compact interface for retrieval, ranking, and generative recommendation. However, effective SID learning is hindered by collisions, where different items are assigned identical or highly confusable codes. Existing methods mainly rely on improved quantization or fixed overlap regularization, but they do not adaptively distinguish whether an overlap should be suppressed or preserved. We propose AdaSID, an adaptive semantic ID learning framework for recommendation. AdaSID regulates SID overlaps through a two-stage process. First, it relaxes repulsion for observed overlaps when the involved items are semantically compatible, preserving admissible sharing rather than uniformly separating all collisions. Second, it allocates the remaining regulation pressure according to local collision load and training progress, strengthening control in congested regions while gradually rebalancing optimization toward recommendation alignment. This design adaptively decides which overlaps to penalize, how strongly to regulate them, and when to shift the learning focus. Extensive offline and online experiments validate AdaSID. On two public benchmarks, AdaSID improves Recall and NDCG by about 4.5% on average over strong baselines, while improving codebook utilization and SID diversity. In Kuaishou e-commerce, an online A/B test on short-video retrieval covering tens of millions of users achieves statistically significant gains, including a 0.98% GMV improvement, and industrial ranking evaluation shows consistent AUC improvements.

推荐系统语义编码多模态自适应学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。