arXiv:2606.06970cs.IR2026-06被引 3

用动态语义ID提升直播推荐,实时反映内容变化并融合用户互动信号。

SSRLive: Live Streaming Recommendation with Dynamic Semantic ID

论文配图:SSRLive: Live Streaming Recommendation with Dynamic Semantic ID
图 1 · 摘自论文原文
  • 设计生成与判别双模块,动态生成语义ID捕捉直播内容变化。
  • 在线测试显示观看时长提升3.38%,商品交易额增长0.72%。
  • 适合关注直播推荐优化与实时交互建模的工业界研究者。

直播已成为增长最快的在线媒体形式之一,支持实时内容传播和用户与主播间的即时互动。尽管现有推荐算法在该领域表现良好,但普遍存在计算资源利用率低、浮点运算量(FLOPs)不足的问题,限制了性能提升。生成式推荐技术虽在多个工业任务中取得进展,但直接应用于直播场景面临两大挑战:(1) 静态语义ID(SIDs)无法反映直播内容的快速变化;(2) 生成流程通常不包含用户-主播交互信号(如点赞、下单),而这些信号对建模用户意图至关重要。为此,我们提出SSRLive:基于动态语义ID的直播推荐框架。该框架在统一架构中融合生成与判别模块:生成部分采用编码器-解码器结构,生成静态与动态SIDs,实现对直播内容的及时表征,并利用多模态信息;判别部分结合SIDs与用户特征,注入用户-主播交互数据,进行多任务预测。真实部署的在线A/B测试表明,该方案带来显著收益:观看时长提升3.38%,商品交易额(GMV)增长0.72%,粉丝增长3.12%,互动量上升2.92%。实验验证了其有效性与商业价值,目前已被全面部署,服务数亿活跃用户。

原文摘要 · Abstract (English)

Live streaming has emerged as one of the fastest-growing forms of online media, enabling instant content broadcasting and real-time engagement between users and streamers. Despite the effectiveness of existing recommendation algorithms in this domain, they often suffer from limited utilization of computational resources, with low FLOPs that hinder further performance enhancement. Generative recommendation techniques, which have gained traction in various industrial tasks, offer a promising avenue for improving live streaming recommendations. However, directly applying generative methods to live streaming is non-trivial due to two major challenges: (1) static semantic IDs (SIDs) cannot reflect the rapidly changing nature of live room content; and (2) generative pipelines generally do not incorporate user--streamer interaction signals (e.g., likes, orders), which are critical for modeling user intent toward both the streamer and showcased products. To address these challenges, we introduce SSRLive: Dynamic Semantic ID-guided Streaming Recommendation for Live platforms. The proposed framework integrates a generative module and a discriminative module in a unified architecture. The generative component employs an encoder-decoder design to produce both static and dynamic SIDs, enabling timely representation of live room content while leveraging multimodal information. The discriminative component refines task-specific representations by combining SIDs with user features, augments them with user-streamer interaction data, and performs multi-task predictions. Online A/B tests in real-world deployment demonstrate tangible benefits: watch time (+3.38%), GMV (+0.72%), follower growth (+3.12%), and interaction volume (+2.92%). These improvements highlight the effectiveness and business value of SSRLive, which is now fully deployed, serving hundreds of millions of active users.

直播推荐动态表征生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。