arXiv:2511.00530cs.IR2025-11NeurIPS被引 15

用扩散模型预测用户未来行为序列,更准更连贯。

Listwise Preference Diffusion Optimization for User Behavior Trajectories Prediction

  • 基于扩散模型直接优化整个行为序列的全局偏好关系。
  • 在真实数据集上,轨迹预测准确率显著超越现有方法。
  • 适合做个性化推荐、内容推送等需要长序列预判的场景。

预测多步用户行为轨迹需要对未来的动作序列进行结构化偏好推理,这是传统序列推荐忽略的问题。该问题对个性化电商和自适应内容分发至关重要,准确预测用户完整行为序列能提升满意度与商业效果。我们发现现有方法的核心缺陷:无法捕捉序列项间的全局列表级依赖。为此,我们提出用户行为轨迹预测(UBTP)新任务,显式建模长期用户偏好。引入列表偏好扩散优化(LPDO),一种基于扩散的训练框架,直接优化整个项目序列的结构化偏好。LPDO融合Plackett-Luce监督信号,并推导出与列表排序似然一致的紧致变分下界,实现去噪步骤间偏好生成的一致性,克服了先前扩散方法的独立标记假设。为严格评估多步预测质量,我们提出任务专用指标序列匹配(SeqMatch),衡量轨迹完全一致度;并采用困惑度(PPL)评估概率保真度。在真实用户行为基准上的大量实验表明,LPDO持续优于当前最先进基线,确立了扩散模型在结构化偏好学习中的新基准。

原文摘要 · Abstract (English)

Forecasting multi-step user behavior trajectories requires reasoning over structured preferences across future actions, a challenge overlooked by traditional sequential recommendation. This problem is critical for applications such as personalized commerce and adaptive content delivery, where anticipating a user's complete action sequence enhances both satisfaction and business outcomes. We identify an essential limitation of existing paradigms: their inability to capture global, listwise dependencies among sequence items. To address this, we formulate User Behavior Trajectory Prediction (UBTP) as a new task setting that explicitly models long-term user preferences. We introduce Listwise Preference Diffusion Optimization (LPDO), a diffusion-based training framework that directly optimizes structured preferences over entire item sequences. LPDO incorporates a Plackett-Luce supervision signal and derives a tight variational lower bound aligned with listwise ranking likelihoods, enabling coherent preference generation across denoising steps and overcoming the independent-token assumption of prior diffusion methods. To rigorously evaluate multi-step prediction quality, we propose the task-specific metric Sequential Match (SeqMatch), which measures exact trajectory agreement, and adopt Perplexity (PPL), which assesses probabilistic fidelity. Extensive experiments on real-world user behavior benchmarks demonstrate that LPDO consistently outperforms state-of-the-art baselines, establishing a new benchmark for structured preference learning with diffusion models.

行为预测扩散模型序列推荐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。