用观察数据优化电商推荐排序策略,提升平台整体收益。
Ranking Policy Learning via Marketplace Expected Value Estimation From Observational Data
- 基于用户浏览行为建模,将排序视为干预动作以最大化预期互动
- 通过上下文价值分布估计平台总收益,考虑会话间差异与数据偏移
- 适用于电商搜索、推荐系统优化,尤其关注收益与用户体验平衡
我们构建了一个决策框架,将电商平台搜索或推荐系统的排序策略学习问题,转化为基于观察数据的期望收益优化问题。作为价值分配机制,排序策略将检索结果分配至指定位置,以在购物旅程的任一阶段最大化用户从展示项中获得的效用。该目标可定义为:在给定排序上下文下,与用户意图匹配的交互事件的期望数量,基于潜在的用户浏览概率模型。通过识别排序作为影响用户与展示项交互的干预行为,并关联其对平台的经济价值,我们将平台的期望收益定义为所有展示排序动作产生的集体价值。公式中的关键要素是上下文价值分布,它不仅体现单次会话内排序干预的价值归属,也刻画了平台收益在不同用户会话间的分布。我们基于观察数据构建了经验性的平台期望收益估计,能处理会话上下文间经济价值异质性以及从用户活动数据学习时的分布偏移。排序策略可通过标准贝叶斯推断技术,以优化经验期望收益估计进行训练。我们在一家大型电商平台的产品搜索任务上报告了实证结果,展示了在不同上下文价值分布极端选择下,由经验收益估计训练出的排序策略所呈现的根本权衡。
原文摘要 · Abstract (English)
We develop a decision making framework to cast the problem of learning a ranking policy for search or recommendation engines in a two-sided e-commerce marketplace as an expected reward optimization problem using observational data. As a value allocation mechanism, the ranking policy allocates retrieved items to the designated slots so as to maximize the user utility from the slotted items, at any given stage of the shopping journey. The objective of this allocation can in turn be defined with respect to the underlying probabilistic user browsing model as the expected number of interaction events on presented items matching the user intent, given the ranking context. Through recognizing the effect of ranking as an intervention action to inform users' interactions with slotted items and the corresponding economic value of the interaction events for the marketplace, we formulate the expected reward of the marketplace as the collective value from all presented ranking actions. The key element in this formulation is a notion of context value distribution, which signifies not only the attribution of value to ranking interventions within a session but also the distribution of marketplace reward across user sessions. We build empirical estimates for the expected reward of the marketplace from observational data that account for the heterogeneity of economic value across session contexts as well as the distribution shifts in learning from observational user activity data. The ranking policy can then be trained by optimizing the empirical expected reward estimates via standard Bayesian inference techniques. We report empirical results for a product search ranking task in a major e-commerce platform demonstrating the fundamental trade-offs governed by ranking polices trained on empirical reward estimates with respect to extreme choices of the context value distribution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。