arXiv:2512.13726cs.LGcs.AI2025-12被引 4

用强化学习优化电商推荐,让有限时间内推荐更高效

Time-Constrained Recommendations: Reinforcement Learning Strategies for E-Commerce

  • 将推荐问题建模为带时间预算的马尔可夫决策过程
  • 在阿里重排数据集上验证,强化学习比传统方法提升效率
  • 适合研究推荐系统资源约束与用户行为的学者

与传统推荐任务不同,用户有限的时间预算构成关键资源约束,要求推荐系统在项目相关性与评估成本间取得平衡。例如,在移动端购物界面中,用户通过滑动浏览推荐列表(slate),每次滑动触发一组项目展示,而用户需花费时间评估项目特征以决定是否点击。高相关性项目若评估成本过高,可能超出用户时间预算,影响参与度。本文旨在评估能同时学习用户偏好与时间预算的强化学习算法,以在资源约束下生成更具参与潜力的推荐。实验基于阿里巴巴个性化重排数据集,探索强化学习在电商场景下对slate优化的应用。贡献包括:(i) 将时间约束下的slate推荐统一建模为具有预算感知效用的马尔可夫决策过程;(ii) 构建仿真框架用于分析策略在重排数据上的行为;(iii) 实证表明,在紧时间预算下,基于策略的和离策略的控制方法优于传统的上下文赌博机方法。

原文摘要 · Abstract (English)

Unlike traditional recommendation tasks, finite user time budgets introduce a critical resource constraint, requiring the recommender system to balance item relevance and evaluation cost. For example, in a mobile shopping interface, users interact with recommendations by scrolling, where each scroll triggers a list of items called slate. Users incur an evaluation cost - time spent assessing item features before deciding to click. Highly relevant items having higher evaluation costs may not fit within the user's time budget, affecting engagement. In this position paper, our objective is to evaluate reinforcement learning algorithms that learn patterns in user preferences and time budgets simultaneously, crafting recommendations with higher engagement potential under resource constraints. Our experiments explore the use of reinforcement learning to recommend items for users using Alibaba's Personalized Re-ranking dataset supporting slate optimization in e-commerce contexts. Our contributions include (i) a unified formulation of time-constrained slate recommendation modeled as Markov Decision Processes (MDPs) with budget-aware utilities; (ii) a simulation framework to study policy behavior on re-ranking data; and (iii) empirical evidence that on-policy and off-policy control can improve performance under tight time budgets than traditional contextual bandit-based methods.

强化学习推荐系统时间约束电商

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。