提出连续生成式观看时长预测方法,解决传统模型的偏差与延迟问题。
FlowTime: Towards Continuous Generative Watch Time Prediction via Flow-based Personalized Priors

- 基于流模型构建个性化先验,动态捕捉用户行为的多模态特征。
- 在多个数据集上超越现有方法,离线与线上测试均显著提升性能。
- 开源首个观看时长预测库,支持可复现实验与新模型评估。
观看时长已成为短视频推荐系统中优化用户参与度的关键指标。然而,当前预测方法存在固有缺陷:直接回归因单峰高斯假设导致均值坍缩,有序回归受刚性离散化影响产生量化误差,离散生成回归则面临高推理延迟和启发式词表设计问题。更深层问题是无法捕捉用户-物品交互模式的内在多模态性和异质性。为此,本文从因果视角重新审视问题,将用户特定行为模式识别为影响观看时长结果的结构性混杂因子,相同兴趣在不同用户习惯下表现为不同观看时长。我们正式提出第四种范式——连续生成回归,并引入FlowTime方法,采用一步生成变分自编码器,在避免迭代去噪延迟的同时保持连续潜空间表达力。进一步设计基于流的个性化先验(Flow-based Personalized Prior),利用神经网络变换标准高斯先验为复杂、历史条件化的流形,实现对多模态交互模式的自适应建模。最后,我们构建了TimeRec——首个开源观看时长预测库,并提出新型个性化度量以建立严格基准。大量离线实验与在线A/B测试表明,FlowTime显著优于现有最优方法。
原文摘要 · Abstract (English)
Watch time has emerged as a pivotal metric for optimizing deep user engagement in short-video recommender systems. However, current methods of watch time prediction (WTP) suffer from inherent paradigm-specific limitations. Direct Regression faces mean-collapse due to unimodal Gaussian assumptions, while Ordinal Regression is hampered by quantization errors from rigid discretization. Similarly, Discrete Generative Regression struggles with high inference latency and heuristic vocabulary design. Beyond these specific flaws, a shared deficiency is the inability to capture the intrinsic multimodality and heterogeneity of User-Item Interaction Patterns. To address these challenges, we first revisit the WTP problem from a causal perspective and identify these user-specific patterns as structural confounders that modulate watch time outcomes, where identical interests manifest as distinct watch time outcomes conditioned on diverse user habits. Then, we formally propose a new (or the fourth) paradigm -- Continuous Generative Regression, and introduce FlowTime, a novel method utilizing a One-step Generative Variational Autoencoder. FlowTime effectively circumvents the latency of iterative denoising while maintaining the expressivity of continuous latent spaces. Furthermore, we design a Flow-based Personalized Prior that leverages NFs to warp a standard Gaussian prior into a complex, history-conditioned manifold, thereby enabling the adaptive modeling of multimodal interaction patterns. Finally, we build TimeRec, the first open-source WTP Library, alongside a novel personalization metric to establish a rigorous benchmarking standard. Extensive offline experiments and online A/B tests demonstrate FlowTime's significant superiority over SOTA methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。