多臂序列检验中,如何在不知哪个臂有效时仍保持最优检测性能。
Multi-Armed Sequential Hypothesis Testing by Betting
- 设计了一种类似置信上界但适用于不可观测奖励的自适应选臂策略
- 证明了在多个非零臂存在时,仍能实现与已知最优臂相当的拒绝时间
- 提出的新方法对药效试验等场景中的高效决策具有实用价值
我们研究一种多臂序列检验的投注方法:每步可选择一个数据源(臂)获取信息。目标是检验复合零假设 𝒫(所有臂均无效,如所有药物剂量均无效果),并拒绝它以支持复合备择假设 𝒬(至少有一个臂有效)。我们提出一个最优性要求:即使多个臂为非零,其表现也应等同于拥有先验知识、知道哪个臂提供最强反证据的情况。为此,我们将对数最优性和期望拒绝时间最优性推广到多臂情形,给出了匹配的上下界。关键技术是设计一种改进的上界类似算法,用于不可观测但足够‘可估计’的奖励。在此过程中,推导出凯利[1956]意义下最优财富增长速率的非渐近集中不等式,可能具有独立研究价值。
原文摘要 · Abstract (English)
We consider a variant of sequential testing by betting where, at each time step, the statistician is presented with multiple data sources (arms) and obtains data by choosing one of the arms. We consider the composite global null hypothesis $\mathscr{P}$ that all arms are null in a certain sense (e.g. all dosages of a treatment are ineffective) and we are interested in rejecting $\mathscr{P}$ in favor of a composite alternative $\mathscr{Q}$ where at least one arm is non-null (e.g. there exists an effective treatment dosage). We posit an optimality desideratum that we describe informally as follows: even if several arms are non-null, we seek $e$-processes and sequential tests whose performance are as strong as the ones that have oracle knowledge about which arm generates the most evidence against $\mathscr{P}$. Formally, we generalize notions of log-optimality and expected rejection time optimality to more than one arm, obtaining matching lower and upper bounds for both. A key technical device in this optimality analysis is a modified upper-confidence-bound-like algorithm for unobservable but sufficiently "estimable" rewards. In the design of this algorithm, we derive nonasymptotic concentration inequalities for optimal wealth growth rates in the sense of Kelly [1956]. These may be of independent interest.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。