arXiv:2503.11819cs.LGcs.GT2025-03

动态定价与商品组合优化,基于用户上下文实时提升收益。

Online Assortment and Price Optimization Under Contextual Choice Models

  • 结合用户上下文信息,实时调整商品组合与价格。
  • 实现近似最优的收益遗憾,达 $\widetilde{O}(d \sqrt{KT}/L_0)$。
  • 适用于电商、推荐系统等需个性化决策场景。

我们研究一种商品组合与定价问题,卖家有 $N$ 种商品可售。每轮中,卖家观测到用户 $d$ 维上下文偏好向量,从 $K$ 种商品中选择一个组合并设定价格。用户根据参数未知的多项式对数选择模型最多选择一件商品。卖家在每轮结束后观察用户是否购买及具体商品,目标是最大化 $T$ 轮内的累计收益。本文提出一种算法,通过用户反馈学习,实现收益遗憾为 $\widetilde{O}(d \sqrt{KT}/L_0)$,其中 $L_0$ 是最小价格敏感度参数。同时证明任何算法的遗憾下界为 $Ω(d \sqrt{T}/L_0)$。

原文摘要 · Abstract (English)

We consider an assortment selection and pricing problem in which a seller has $N$ different items available for sale. In each round, the seller observes a $d$-dimensional contextual preference information vector for the user, and offers to the user an assortment of $K$ items at prices chosen by the seller. The user selects at most one of the products from the offered assortment according to a multinomial logit choice model whose parameters are unknown. The seller observes which, if any, item is chosen at the end of each round, with the goal of maximizing cumulative revenue over a selling horizon of length $T$. For this problem, we propose an algorithm that learns from user feedback and achieves a revenue regret of order $\widetilde{O}(d \sqrt{K T} / L_0 )$ where $L_0$ is the minimum price sensitivity parameter. We also obtain a lower bound of order $Ω(d \sqrt{T}/ L_0)$ for the regret achievable by any algorithm.

动态定价在线优化上下文选择收益管理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。