arXiv:2509.22807cs.IRcs.AI2025-09NeurIPS被引 1

用心理奖励模型纠正点击误导,让推荐更贴近用户真实喜好。

MTRec: Learning to Align with User Preferences via Mental Reward Models

  • 通过心理奖励模型量化用户对推荐内容的真实满意度
  • 在工业级短视频平台实现平均观看时长提升7%
  • 适合需要精准理解用户偏好的推荐系统优化场景

推荐系统主要依赖隐式用户反馈(如点击)进行训练,但这类信号常不能反映用户真实偏好——例如用户因标题吸引点击了新闻,阅读后却感到不适。缺乏显式反馈时,此类错误信号会严重误导推荐系统。本文提出MTRec,一种新型序列推荐框架,通过揭示用户对推荐内容的内在满意度来对齐真实偏好。具体地,引入心理奖励模型量化用户满意度,并采用分布式逆强化学习方法进行学习。所学心理奖励模型用于引导推荐模型更好地匹配用户真实偏好。实验表明,MTRec显著提升多种推荐模型性能;在工业级短视频平台部署后,平均用户观看时长提升7%。

原文摘要 · Abstract (English)

Recommendation models are predominantly trained using implicit user feedback, since explicit feedback is often costly to obtain. However, implicit feedback, such as clicks, does not always reflect users' real preferences. For example, a user might click on a news article because of its attractive headline, but end up feeling uncomfortable after reading the content. In the absence of explicit feedback, such erroneous implicit signals may severely mislead recommender systems. In this paper, we propose MTRec, a novel sequential recommendation framework designed to align with real user preferences by uncovering their internal satisfaction on recommended items. Specifically, we introduce a mental reward model to quantify user satisfaction and propose a distributional inverse reinforcement learning approach to learn it. The learned mental reward model is then used to guide recommendation models to better align with users' real preferences. Our experiments show that MTRec brings significant improvements to a variety of recommendation models. We also deploy MTRec on an industrial short video platform and observe a 7 percent increase in average user viewing time.

推荐系统心理建模逆强化学习用户偏好

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。