让推荐过程实时推理,提升复杂偏好捕捉能力
Reasoning While Recommending: Entropy-Guided Latent Reasoning in Generative Re-ranking Models
- 边生成推荐列表边实时推理,打破先推理后推荐的旧模式
- 通过动态温度调节和上下文感知推理,降低决策熵,提升精度
- 轻量集成设计,可无缝增强现有生成式重排序模型
强化学习在生成式重排序中至关重要,因其具备探索与利用的平衡能力。然而,现有方法难以适应生成过程中模型难度带来的动态熵变,导致难以准确捕捉复杂偏好。受语言模型通过推理实现突破的启发,本文提出一种潜在推理机制,实验证明该机制有效降低决策熵。基于此,我们构建了熵引导的潜在推理(EGLR)推荐模型,具有三大优势:一、摒弃“先推理后推荐”范式,实现“边推荐边推理”,专为高难度列表生成设计;二、采用上下文感知推理标记与动态温度调节,实现可变长度推理,扩大探索范围,提升推荐精度,优化探索-利用权衡;三、采用轻量级集成设计,无需复杂独立模块或后处理,易于适配现有模型。在两个真实数据集上的实验验证了模型有效性,其显著优势在于可兼容现有生成式重排序模型以提升性能。进一步分析也展示了其部署价值与研究潜力。
原文摘要 · Abstract (English)
Reinforcement learning plays a crucial role in generative re-ranking scenarios due to its exploration-exploitation capabilities, but existing generative methods mostly fail to adapt to the dynamic entropy changes in model difficulty during list generation, making it challenging to accurately capture complex preferences. Given that language models have achieved remarkable breakthroughs by integrating reasoning capabilities, we draw on this approach to introduce a latent reasoning mechanism, and experimental validation demonstrates that this mechanism effectively reduces entropy in the model's decision-making process. Based on these findings, we introduce the Entropy-Guided Latent Reasoning (EGLR) recommendation model, which has three core advantages. First, it abandons the "reason first, recommend later" paradigm to achieve "reasoning while recommending", specifically designed for the high-difficulty nature of list generation by enabling real-time reasoning during generation. Second, it implements entropy-guided variable-length reasoning using context-aware reasoning token alongside dynamic temperature adjustment, expanding exploration breadth in reasoning and boosting exploitation precision in recommending to achieve a more precisely adapted exploration-exploitation trade-off. Third, the model adopts a lightweight integration design with no complex independent modules or post-processing, enabling easy adaptation to existing models. Experimental results on two real-world datasets validate the model's effectiveness, and its notable advantage lies in being compatible with existing generative re-ranking models to enhance their performance. Further analyses also demonstrate its practical deployment value and research potential.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。