arXiv:2607.14706cs.LG2026-07

提出抗策略干扰的线性多臂老虎机最佳臂识别算法

MESHA: Mechanism-Enforced Sequential Halving for Strategic Linear Bandits

  • 用均匀采样降低策略行为影响,结合分轮惩罚机制剔除异常报告
  • 在固定预算T内,失败概率有上界,可保证最优臂被正确识别
  • 适合对抗恶意伪装特征的场景,对传统方法有显著优势

我们设计并分析了机制强化的分阶段减半算法(MESHA),用于战略线性多臂老虎机中的最佳臂识别(BAI)。在此设定中,每个臂可能通过虚假报告其特征向量来最大化被识别为最优臂的概率,而奖励由真实但不可观测的特征生成。MESHA采用朴素均匀采样规则和分轮格里姆触发条件(GTC):前者减少策略行为的影响,后者剔除报告特征严重偏离真实值的臂。针对任意纳什均衡,我们证明任何臂都会试图通过GTC检查以最大化被识别概率,并推导出在固定预算T下的失败概率上界。我们还表明,基于G-最优设计的先进线性BAI算法在此类策略环境中会失效,因为基于虚假报告特征的最优设计采样规则可能导致最优臂完全无法获得采样预算。大量数值实验表明,MESHA优于依赖最优设计采样规则及特征无关基线的方法,验证了其有效性。

原文摘要 · Abstract (English)

We design and analyze \underline{M}echanism-\underline{E}nforced \underline{S}equential \underline{HA}lving (MESHA), an algorithm for Best Arm Identification (BAI) in strategic linear bandits. In this setting, each arm may strategically misreport its feature vector to maximize the probability of being identified as the best arm, when rewards are generated from the arms' true but unobservable features. The design of MESHA applies the naïve uniform sampling rule and an epoch-wise Grim Trigger Condition (GTC): the former reduces the impact of arms' strategic behaviours and the latter eliminates arms whose reported features severely deviate from the ground truth. Considering an arbitrary Nash Equilibrium, we prove that any arm would attempt to pass the GTC check to maximize its identified probability and derive an upper bound on the failure probability of MESHA within a fixed budget $T$. We also show that state-of-the-art linear BAI algorithms with $G$-optimal design would fail in such strategic environment, as the optimal design (OD)-based sampling rule based on strategically reported features may {\it starve} the optimal arm of any sampling budget. Finally, extensive numerical experiments indicate that MESHA outperforms baselines that rely on OD-based sampling rules as well as the feature-agnostic baselines, corroborating the efficacy of MESHA.

多臂老虎机机制设计策略学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。