平台如何通过选品策略高效学习新产品的质量,平衡探索与收益。
Optimal Exploration of New Products under Assortment Decisions
- 新商品需搭配头部老商品一同展示以提升学习效率。
- 同时探索多个新商品时,数量由其潜力决定而非销量概率。
- 经典算法在该场景下表现不佳,需定制化策略。
我们研究平台在容量受限的选品决策下,对新产品的在线学习问题。新商品初始质量未知,通过用户购买并留下评价实现社会性知识传播。由于评价依赖购买行为,平台必须将新商品纳入选品(即探索)以获取反馈,但此举成本较高,因新商品需求低于成熟商品。本文刻画了最小化遗憾的最优选品策略,回答两个关键问题:(1) 新商品应单独展示还是与成熟商品组合?尽管单独展示可提高购买概率,但研究发现始终最优是将其与最高排名的成熟商品配对。(2) 当存在多个新商品时,应同时探索还是逐一探索?我们证明最优同时探索数量具有简单阈值结构:随新商品潜力上升而增加,且不依赖于其个体购买概率。此外,两种经典强化学习算法——UCB与Thompson Sampling——在此场景中均失败:前者过度探索,后者探索不足。研究结果为平台如何通过选品策略学习新产品提供了结构性洞见。
原文摘要 · Abstract (English)
We study online learning for new products on a platform that makes capacity-constrained assortment decisions on which products to offer. For a newly listed product, its quality is initially unknown, and quality information propagates through social learning: when a customer purchases a new product and leaves a review, its quality is revealed to both the platform and future customers. Since reviews require purchases, the platform must feature new products in the assortment ("explore") to generate reviews to learn about new products. Such exploration is costly because customer demand for new products is lower than for incumbent products. We characterize the optimal assortments for exploration to minimize regret, addressing two questions. (1) Should the platform offer a new product alone or alongside incumbent products? The former maximizes the purchase probability of the new product but yields lower short-term revenue. Despite the lower purchase probability, we show it is always optimal to pair the new product with the top incumbent products. (2) With multiple new products, should the platform explore them simultaneously or one at a time? We show that the optimal number of new products to explore simultaneously has a simple threshold structure: it increases with the "potential" of the new products and, surprisingly, does not depend on their individual purchase probabilities. We also show that two canonical bandit algorithms, UCB and Thompson Sampling, both fail in this setting for opposite reasons: UCB over-explores while Thompson Sampling under-explores. Our results provide structural insights on how platforms should learn about new products through assortment decisions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。