arXiv:2504.08771cs.IRcs.AI2025-04

通过模拟用户观看过程预测短视频停留时长,提升推荐效果。

Generate the browsing process for short-video recommendation

  • 用用户互动行为建模观看旅程,隐式捕捉兴趣变化。
  • 在工业级数据上实现83% XAUC,APP使用时长提升0.13%。
  • 适合需要精准时长预测的短视频推荐系统应用。

本文提出一种生成方法,动态模拟用户在短视频推荐中的观看过程,以预测观看时长。与依赖多模态特征理解视频内容的现有方法不同,该方法通过学习协同信息,利用正负反馈视频的兴趣变化和用户交互行为,隐式建模用户的观看旅程。通过基于时长划分视频片段并采用类似Transformer的架构,有效捕捉片段间的序列依赖关系,同时缓解时长偏差问题。在工业级和公开数据集上的大量实验表明,该方法在观看时长预测任务中达到领先性能。该方法已部署于快手极速版,在大规模流式训练集上实现单视频观看时长预测的XAUC达83%,较其他方法显著提升,且使APP使用时长提高0.13%。所提方法通过片段级建模与用户参与反馈,为视频推荐提供了可扩展且高效的解决方案。

原文摘要 · Abstract (English)

This paper proposes a generative method to dynamically simulate users' short video watching journey for watch time prediction in short video recommendation. Unlike existing methods that rely on multimodal features for video content understanding, our method simulates users' sustained interest in watching short videos by learning collaborative information, using interest changes from existing positive and negative feedback videos and user interaction behaviors to implicitly model users' video watching journey. By segmenting videos based on duration and adopting a Transformer-like architecture, our method can capture sequential dependencies between segments while mitigating duration bias. Extensive experiments on industrial-scale and public datasets demonstrate that our method achieves state-of-the-art performance on watch time prediction tasks. The method has been deployed on Kuaishou Lite, achieving a significant improvement of +0.13\% in APP duration, and reaching an XAUC of 83\% for single video watch time prediction on industrial-scale streaming training sets, far exceeding other methods. The proposed method provides a scalable and effective solution for video recommendation through segment-level modeling and user engagement feedback.

短视频推荐观看时长预测用户行为建模序列建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。