考虑顾客到访受商品组合和价格影响的动态定价与选品策略。
Poisson-MNL Bandit: Nearly Optimal Dynamic Joint Assortment and Pricing with Decision-Dependent Customer Arrivals
- 将顾客选择与到访率耦合建模,用泊松-多分类逻辑斯蒂模型描述依赖决策的客流。
- 提出PMNL算法,实现近最优的累积收益,理论证明其后悔率约为√(T log T)。
- 适合研究动态定价、零售优化或需考虑客流反馈的场景,尤其在非固定客流量下表现优异。
我们研究动态联合选品与定价问题,卖家在固定周期内更新决策以最大化周期内累计收益。在许多场景中,提供的商品组合与价格不仅影响顾客购买行为,还影响期内到访顾客数量;而经典多分类逻辑斯蒂(MNL)模型假设到访数固定,可能导致次优决策。为此,我们提出泊松-MNL模型,将上下文感知的MNL选择模型与依赖于商品组合和价格的泊松到达模型相结合。基于该模型,我们设计了高效算法PMNL,采用上置信界(UCB)思想。理论上,我们证明了其非渐近后悔界为O(√(T log T)),并给出了匹配的下界(相差log T因子)。模拟实验表明,考虑到达率对决策的依赖至关重要:PMNL能有效学习顾客选择与到达模型,生成优于假设固定到达率的联合选品与定价策略。
原文摘要 · Abstract (English)
We study dynamic joint assortment and pricing where a seller updates decisions at regular accounting/operating intervals to maximize the cumulative per-period revenue over a horizon $T$. In many settings, assortment and prices affect not only what an arriving customer buys but also how many customers arrive within the period, whereas classical multinomial logit (MNL) models assume arrivals as fixed, potentially leading to suboptimal decisions. We propose a Poisson-MNL model that couples a contextual MNL choice model with a Poisson arrival model whose rate depends on the offered assortment and prices. Building on this model, we develop an efficient algorithm PMNL based on the idea of upper confidence bound (UCB). We establish its (near) optimality by proving a non-asymptotic regret bound of order $\sqrt{T\log{T}}$ and a matching lower bound (up to $\log T$). Simulation studies underscore the importance of accounting for the dependency of arrival rates on assortment and pricing: PMNL effectively learns customer choice and arrival models and provides joint assortment-pricing decisions that outperform others that assume fixed arrival rates.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。