arXiv:2605.17238cs.LGstat.ML2026-05

优化商品推荐位置与组合,让平台收益更高

Learning in Position-Aware Multinomial Logit Bandits: From Multiplicative to General Position Effects

论文配图:Learning in Position-Aware Multinomial Logit Bandits: From Multiplicative to General Position Effects
图 1 · 摘自论文原文
  • 设计轮次更新算法,实时调整商品展示位置和组合
  • 在乘法型位置效应下实现最优后悔率,比之前提升√K倍
  • 适用于电商平台动态推荐,对算法和运营都实用

我们研究动态联合商品组合与位置选择问题,在多项式逻辑(MNL)选择框架下,每个商品的吸引力同时取决于其自身属性和展示位置。研究从乘法型位置效应模型(商品吸引力乘以位置系数)扩展到一般位置效应模型(每种商品-位置组合有独立吸引力参数),以捕捉异质协同效应。针对两类模型,我们设计了基于轮次的在线学习算法,可在每次反馈后立即更新策略,并首次建立了后悔率最优的理论刻画。对于乘法模型,提出带裁剪机制的跨位置成对最大似然估计器,证明算法P2MLE-UCB达到$ ilde{O}( ext{sqrt}{NT})$的后悔率,匹配下界,解决了以往分段分析遗留的$ ext{sqrt}{K}$差距。对于一般模型,建立极小极大下界,并提出GP2-UCB算法实现匹配上界。此外,基于Dinkelbach方法和最大权二分图匹配,设计高效每轮优化子程序。在合成数据和Expedia数据集上的实验表明,所提算法持续优于现有最优基准。

原文摘要 · Abstract (English)

We study the dynamic joint assortment selection and positioning problem, where the attraction of each product depends on both its intrinsic appeal and its display position under a Multinomial Logit (MNL) choice framework. Our study ranges from the multiplicative position effects model, in which each product's attraction is scaled by a position-specific factor, to a general position effects model assigning independent attraction parameters to every product--position pair to capture heterogeneous synergies. For both models, we design round-based learning algorithms that update decisions after every single feedback, and establish the first regret-optimal characterization. Besides, our round-based algorithms provide the prompt operations needed by modern platforms. For the multiplicative model, we develop a cross-position pairwise maximum likelihood estimator with a clipping mechanism, and prove that our algorithm P2MLE-UCB attains a regret of $\tilde{O}(\sqrt{NT})$, matching the lower bound and closing the $\sqrt{K}$ gap left by prior epoch-based analyses. For the general model, we establish a minimax lower bound and propose GP2-UCB with a matching upper bound. Moreover, we design an efficient subroutine for the per-round joint assortment and positioning optimization based on Dinkelbach's method and maximum-weight bipartite matching. Numerical experiments on synthetic data and the Expedia dataset show that our algorithms consistently outperform state-of-the-art benchmarks.

推荐系统在线学习多臂赌博机动态优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。