arXiv:2506.02261cs.IRcs.LG2025-06ACL被引 1

LLM推荐效果关键在感知偏好强弱与时间上下文。

What Makes LLMs Effective Sequential Recommenders? A Study on Preference Intensity and Temporal Context

  • 将显式隐式反馈统一为结构化偏好信号
  • 自适应奖励边界提升跨数据集性能
  • 适合关注用户行为动态建模的研究者

为何大语言模型能有效建模用户偏好?研究发现,现有对齐方法主要依赖二元比较,忽略了两个关键因素:偏好强度(亲疏程度的结构性)和时间上下文(近期行为更能反映当前意图)。通过受控实验表明,利用包含完整反馈的结构化偏好信号可显著提升推荐性能,说明二元建模会丢失重要信息。为此提出RecPO框架,将显式与隐式反馈映射为统一偏好信号,并构建兼顾偏好强度与交互时效性的自适应奖励边界。在五个数据集上的实验显示,RecPO持续优于现有最优基线,且表现出符合人类决策的行为模式,包括偏好即时满足、维持偏好一致性、避免不偏好项。结果表明,偏好强度与时间上下文是基于LLM推荐的有效核心要素。

原文摘要 · Abstract (English)

What enables large language models (LLMs) to effectively model user preferences in sequential recommendation? Our investigation reveals that existing preference-alignment approaches largely rely on binary pairwise comparisons, overlooking two critical factors: preference intensity (the structured strength of affinity or aversion) and temporal context (the extent to which recent interactions better reflect a user's current intent). Through controlled experiments, we show that leveraging comprehensive feedback with structured preference signals substantially improves recommendation performance, indicating that binary modeling discards essential information. Motivated by these findings, we propose RecPO, a unified preference optimization framework that maps both explicit and implicit feedback into a common preference signal and constructs adaptive reward margins that jointly account for preference intensity and interaction recency. Experiments across five datasets show that RecPO consistently outperforms state-of-the-art baselines while exhibiting behavioral patterns aligned with human decision-making, including favoring immediate satisfaction, maintaining preference coherence, and avoiding dispreferred items. Our results highlight that preference intensity and temporal context are fundamental ingredients for effective LLM-based recommendation.

LLM推荐偏好建模时序建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。