arXiv:2503.00565stat.MLcs.LG2025-03被引 2

提出新方法解决带上下文的分批多臂老虎机问题,提升决策效率。

Batched Single-Index Global Multi-Armed Bandits with Covariates

  • 用单指标回归建模各臂奖励关系,兼顾可解释与灵活
  • 在固定臂数下达到最优理论误差率,避免维度灾难
  • 适合医疗推荐等需高效反馈的实时决策场景

多臂老虎机(MAB)是序列决策中的常用框架,决策者每轮选择一个臂以最大化长期收益。在个性化医疗、推荐系统等实际应用中,决策时有上下文信息,不同臂的奖励相关而非独立,且反馈以批次形式提供。本文提出一种新的半参数分批带上下文的老虎机框架,引入跨臂共享参数。利用单指标回归(SIR)模型捕捉臂间奖励关系,在可解释性与灵活性间取得平衡。所提算法BIDS采用分批逐次淘汰策略,结合由单指标方向引导的动态分箱机制。考虑两种情形:一为已有预估方向,二为从数据中估计方向,分别推导理论后悔界。当预估方向足够准确且臂数K固定时,本方法在d=1情况下达到非参数分批老虎机的极小最大值最优率,克服维度灾难。在模拟和真实数据集上的大量实验表明,其性能优于 extcite{jiang2025batched}提出的非参数分批老虎机方法。

原文摘要 · Abstract (English)

The multi-armed bandits (MAB) framework is a widely used approach for sequential decision-making, where a decision-maker selects an arm in each round with the goal of maximizing long-term rewards. In many practical applications, such as personalized medicine and recommendation systems, contextual information is available at the time of decision-making, rewards from different arms are related rather than independent, and feedback is provided in batches. We propose a novel semi-parametric framework for batched bandits with covariates that incorporates a shared parameter across arms. We leverage the single-index regression (SIR) model to capture relationships between arm rewards while balancing interpretability and flexibility. Our algorithm, Batched single-Index Dynamic binning and Successive arm elimination (BIDS), employs a batched successive arm elimination strategy with a dynamic binning mechanism guided by the single-index direction. We consider two settings: one where a pilot direction is available and another where the direction is estimated from data, deriving theoretical regret bounds for both cases. When a pilot direction is available with sufficient accuracy and the number of arms $K$ is fixed, our approach achieves minimax-optimal rates (with $d = 1$) for nonparametric batched bandits, circumventing the curse of dimensionality. Extensive experiments on simulated and real-world datasets demonstrate the effectiveness of our algorithm compared to the nonparametric batched bandit method introduced by \cite{jiang2025batched}.

多臂老虎机上下文决策分批学习单指标模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。