提出OBSI算法,按置信度逐步引入特征,提升稀疏上下文下的决策公平性与效率。
Batched Online Contextual Sparse Bandits with Sequential Inclusion of Features
- OBSI算法逐次加入高置信度特征,避免无关特征干扰决策
- 在合成数据上,相比其他方法更优的累积损失、特征相关性与计算效率
- 适合关注个性化推荐中特征选择与公平性的研究者
多臂老虎机(MAB)在在线平台和电子商务中广泛用于优化个性化用户体验。本文研究线性奖励下的上下文老虎机问题,在稀疏特征与分批数据条件下,通过一种新算法Online Batched Sequential Inclusion(OBSI)实现决策过程中的公平性:仅当对特征影响回报的置信度足够高时才逐步引入。实验在合成数据上显示,OBSI在累积遗憾、所用特征的相关性以及计算开销方面均优于现有方法。
原文摘要 · Abstract (English)
Multi-armed Bandits (MABs) are increasingly employed in online platforms and e-commerce to optimize decision making for personalized user experiences. In this work, we focus on the Contextual Bandit problem with linear rewards, under conditions of sparsity and batched data. We address the challenge of fairness by excluding irrelevant features from decision-making processes using a novel algorithm, Online Batched Sequential Inclusion (OBSI), which sequentially includes features as confidence in their impact on the reward increases. Our experiments on synthetic data show the superior performance of OBSI compared to other algorithms in terms of regret, relevance of features used, and compute.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。