动态调整商品组合与价格,从用户购买反馈中学习偏好并提升收益
Dynamic Assortment Selection and Pricing with Censored Preference Feedback
- 基于被截断的多项对数选择模型,模拟用户过滤高价商品后择一购买的行为
- 提出新型上下文置信界策略,在3000轮模拟中实现近似最优的累计收益
- 适合研究动态定价、个性化推荐的从业者或算法优化研究人员
本文研究动态多产品选择与定价问题,提出基于被截断多项对数(C-MNL)选择模型的新框架。卖家展示一组带价商品,买家先剔除高于自身估值的商品,再从剩余选项中按偏好最多购买一件。目标是在学习产品估值和用户偏好的同时,通过动态调整商品组合与价格来最大化收益。为此,我们设计了下置信界(LCB)定价策略,并结合上置信界(UCB)或汤普森采样(TS)进行产品选择。理论分析表明,两种算法分别达到$ ilde{O}(d^{rac{3}{2}}\ \sqrt{T/κ})$和$ ilde{O}(d^{2}\ \sqrt{T/κ})$的后悔上界。通过仿真验证,方法在多种场景下表现优异。
原文摘要 · Abstract (English)
In this study, we investigate the problem of dynamic multi-product selection and pricing by introducing a novel framework based on a \textit{censored multinomial logit} (C-MNL) choice model. In this model, sellers present a set of products with prices, and buyers filter out products priced above their valuation, purchasing at most one product from the remaining options based on their preferences. The goal is to maximize seller revenue by dynamically adjusting product offerings and prices, while learning both product valuations and buyer preferences through purchase feedback. To achieve this, we propose a Lower Confidence Bound (LCB) pricing strategy. By combining this pricing strategy with either an Upper Confidence Bound (UCB) or Thompson Sampling (TS) product selection approach, our algorithms achieve regret bounds of $\tilde{O}(d^{\frac{3}{2}}\sqrt{T/κ})$ and $\tilde{O}(d^{2}\sqrt{T/κ})$, respectively. Finally, we validate the performance of our methods through simulations, demonstrating their effectiveness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。