arXiv:2608.24034cs.IR2026-08

为直播广告设计动态生成推荐框架,提升实时转化效果。

TAGR: Temporally Adaptive Generative Recommendation for Industrial Live-Streaming Advertising

论文配图:TAGR: Temporally Adaptive Generative Recommendation for Industrial Live-Streaming Advertising
图 1 · 摘自论文原文
  • 用动态广告标识跟踪直播变化,保持生成稳定性
  • 多粒度建模用户意图,结合业务价值加权预测
  • 周期性更新偏好并兼顾训练稳定,适合工业级应用

直播广告是短视频与电商平台的重要变现渠道,其快速变化的直播内容、推广商品和用户反馈对推荐模型的时效性提出高要求。现有生成式推荐模型在静态场景下表现不佳,主要体现在:静态语义标识无法追踪动态直播广告;单尺度行为建模遗漏意图转移;新鲜在线反馈与训练稳定性之间存在冲突。本文提出TAGR框架,在广告标记、用户意图建模和偏好对齐三个层面实现时间自适应:在标记层面,基于当前直播场景与推广商品定期更新活跃广告的语义协同标识(LSID),同时保留稳定的分层词汇用于自回归生成;在意图层面,将多粒度直播间进入历史作为主意图序列,辅助行为作为独立输入,通过请求后意图证据和业务价值加权下一词预测(NTP);在对齐层面,周期性从当前策略中采样新候选组,交替执行行为与价值对齐的偏好更新及监督式NTP维护,以保持学习到的行为分布。在大规模电商直播广告平台部署后,TAGR使直播间进入率和购物车点击率分别提升8.5%和7.4%,收入相较生产基线提升16.1%,验证了时间自适应生成推荐在直播广告中的有效性和工业可行性。

原文摘要 · Abstract (English)

Live-streaming advertising is an important monetization channel on short-video and e-commerce platforms, where rapidly changing live content, promoted products, and user feedback impose strong freshness requirements on recommendation models. Existing generative recommenders designed for static domains fail at three levels: static semantic IDs (SID) cannot track evolving live ads; single-scale behavior modeling misses shifting intent; preference optimization conflicts between fresh on-policy feedback and training stability. We propose TAGR, a generative recommendation framework with temporal adaptation at three levels: live-ad tokenization, user intent modeling, and preference alignment. At the token level, Live Semantic-Collaborative ID (LSID) periodically refreshes each active ad's SID based on its current live scene and promoted products, while retaining a stable hierarchical token vocabulary for autoregressive generation. At the intent level, Intent-Aware Generation (IAG) models live-room entry histories at multiple temporal granularities as the primary intent sequence, keeps auxiliary behaviors as separate inputs, and weights next-token prediction (NTP) using post-request intent evidence and business value. At the alignment level, Intermittent On-Policy Preference Optimization (IOPO) periodically samples fresh candidate groups from the current policy and performs behavior- and value-aligned preference updates interleaved with supervised NTP maintenance to preserve learned behavior distribution. Deployed on a large-scale e-commerce live-stream advertising platform, TAGR improves live-room entry and shopping-cart click rates by 8.5% and 7.4%, respectively, and achieves a 16.1% revenue lift over the production baseline. These results demonstrate the effectiveness and industrial viability of temporally adaptive generative recommendation for live-stream advertising.

生成推荐直播广告动态建模工业落地

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。