arXiv:2601.03153cs.IR2026-01被引 12

通过并行推理流提升用户行为序列推荐效果,兼顾精度与实时性

Parallel Latent Reasoning for Sequential Recommendation

  • 引入并行潜空间推理,同时探索多条不同路径
  • 在三个真实数据集上超越现有最优方法,且推理不延迟
  • 适合需要高精度推荐的电商、内容平台场景

从稀疏的行为序列中捕捉复杂的用户偏好,仍是顺序推荐中的核心挑战。近年来的潜空间推理方法通过测试时的多步推理扩展计算能力,但仅依赖单一路径的深度扩展,随着推理深度增加收益递减。为此,我们提出全新的并行潜空间推理(PLR)框架,首次实现宽度级计算扩展:通过连续潜空间中的可学习触发标记构建多条并行推理流,利用全局推理正则化保持流间多样性,并通过混合推理流聚合机制自适应融合多路输出。在三个真实世界数据集上的大量实验表明,PLR显著优于当前最优基线,同时保持实时推理效率。理论分析进一步验证了并行推理能有效提升泛化能力。本工作为顺序推荐中的推理能力拓展开辟了新路径。

原文摘要 · Abstract (English)

Capturing complex user preferences from sparse behavioral sequences remains a fundamental challenge in sequential recommendation. Recent latent reasoning methods have shown promise by extending test-time computation through multi-step reasoning, yet they exclusively rely on depth-level scaling along a single trajectory, suffering from diminishing returns as reasoning depth increases. To address this limitation, we propose \textbf{Parallel Latent Reasoning (PLR)}, a novel framework that pioneers width-level computational scaling by exploring multiple diverse reasoning trajectories simultaneously. PLR constructs parallel reasoning streams through learnable trigger tokens in continuous latent space, preserves diversity across streams via global reasoning regularization, and adaptively synthesizes multi-stream outputs through mixture-of-reasoning-streams aggregation. Extensive experiments on three real-world datasets demonstrate that PLR substantially outperforms state-of-the-art baselines while maintaining real-time inference efficiency. Theoretical analysis further validates the effectiveness of parallel reasoning in improving generalization capability. Our work opens new avenues for enhancing reasoning capacity in sequential recommendation beyond existing depth scaling.

顺序推荐并行推理潜空间建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。