arXiv:2604.25834cs.AIcs.IR2026-04

用动作时间建模用户意图,提升短视频推荐精准度

Action-Aware Generative Sequence Modeling for Short Video Recommendation

论文配图:Action-Aware Generative Sequence Modeling for Short Video Recommendation
图 1 · 摘自论文原文
  • 按用户操作时间序列建模,捕捉动态兴趣变化
  • 在线测试中提升观看时长0.34%、互动率8.1%、留存0.162%
  • 适合做短视频平台个性化推荐系统的工程师和研究员

随着互联网快速发展,用户对内容推荐精度要求越来越高。短视频通常包含多样片段,用户对其态度各异,传统将视频视为整体的二分类推荐模型难以捕捉这种细微偏好。本文通过统计分析与行为模式观察发现,用户动作的时间分布能反映多样化意图。基于此,提出新型建模框架A2Gen:在时间维度上细化用户行为并形成序列,统一处理预测。首先引入上下文感知注意力模块(CAM)融合物品特征建模行为序列;进而设计分层序列编码器(HSE)学习历史行为的时间模式;最后利用CAM构建自回归生成模块(AAG)生成行为序列。在快手数据集和天猫公开数据集上的离线实验验证了模型优势;在快手平台的大规模线上A/B测试显示,该模型在多任务预测中显著优于基线方法,分别带来0.34%的观看时长提升、8.1%的互动率增长和0.162%的7日留存率提升,已成功全量部署,服务每日超4亿用户。

原文摘要 · Abstract (English)

With the rapid development of the Internet, users have increasingly higher expectations for the recommendation accuracy of online content consumption platforms. However, short videos often contain diverse segments, and users may not hold the same attitude toward all of them. Traditional binary-classification recommendation models, which treat a video as a single holistic entity, face limitations in accurately capturing such nuanced preferences. Considering that user consumption is a temporal process, this paper demonstrates that the timing of user actions can represent diverse intentions through statistical analysis and examination of action patterns. Based on this insight, we propose a novel modeling paradigm: Action-Aware Generative Sequence Network (A2Gen), which refines user actions along the temporal dimension and chains them into sequences for unified processing and prediction. First, we introduce the Context-aware Attention Module (CAM) to model action sequences enriched with item-specific contextual features. Building upon this, we develop the Hierarchical Sequence Encoder (HSE) to learn temporal action patterns from users' historical actions. Finally, through leveraging CAM, we design a module for action sequence generation: the Action-seq Autoregressive Generator (AAG). Extensive offline experiments on the Kuaishou's dataset and the Tmall public dataset demonstrate the superiority of our proposed model. Furthermore, through large-scale online A/B testing deployed on Kuaishou's platform, our model achieves significant improvements over baseline methods in multi-task prediction by leveraging sequential information. Specifically, it yields increases of 0.34% in user watch time, 8.1% in interaction rate, and 0.162% in overall user retention (LifeTime-7), leading to successful deployment across all traffic, serving over 400 million users every day.

短视频推荐序列建模行为分析A/B测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。