用语言锚点增强直播推荐,让内容理解更精准
SARM: LLM-Augmented Semantic Anchor for End-to-End Live-Streaming Ranking
- 用可学习的文本锚点直接融入排序优化,提升语义匹配精度
- 在离线与大规模线上测试中均优于现有基线,日均服务超4亿用户
- 适合需要高精度实时推荐的直播平台,尤其关注内容理解
大规模直播推荐需在严格实时约束下精确建模动态内容语义。工业部署中,离散语义抽象因聚类损失描述精度,而密集多模态嵌入独立提取且与排序优化弱对齐,限制了细粒度内容感知。为此,我们提出端到端排序架构SARM,将自然语言语义锚点直接融入排序优化,实现基于多模态内容的细粒度作者表征。每个语义锚点以可学习文本标记形式存在,与排序特征联合优化,使模型能根据排序目标自适应内容描述。轻量级双标记门控设计捕捉直播领域特定语义,非对称部署策略保障低延迟在线训练与服务。大量离线评估与大规模A/B测试显示持续优于生产基线。SARM已全量部署,日均服务超4亿用户。
原文摘要 · Abstract (English)
Large-scale live-streaming recommendation requires precise modeling of non-stationary content semantics under strict real-time serving constraints. In industrial deployment, two common approaches exhibit fundamental limitations: discrete semantic abstractions sacrifice descriptive precision through clustering, while dense multimodal embeddings are extracted independently and remain weakly aligned with ranking optimization, limiting fine-grained content-aware ranking. To address these limitations, we propose \textbf{SARM}, an end-to-end ranking architecture that integrates natural-language semantic anchors directly into ranking optimization, enabling fine-grained author representations conditioned on multimodal content. Each semantic anchor is represented as learnable text tokens jointly optimized with ranking features, allowing the model to adapt content descriptions to ranking objectives. A lightweight dual-token gated design captures domain-specific live-streaming semantics, while an asymmetric deployment strategy preserves low-latency online training and serving. Extensive offline evaluation and large-scale A/B tests show consistent improvements over production baselines. SARM is fully deployed and serves over 400 million users daily.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。