用强化学习设计可随时验证的测试方法,提升有限时间内统计推断效率。
Learning to Bet for Horizon-Aware Anytime-Valid Testing
- 基于赌局框架将时序测试建模为最优控制问题,状态由时间与赌注累积值决定。
- 在不同进度下选择不同策略:落后时激进下注,领先时保守下注,效果优于固定策略。
- 通过深度强化学习训练通用智能体,适配多种场景和置信水平,适合实时决策应用。
我们针对严格截止时间 $N$ 下有界均值的统计推断,开发了具备时域感知能力的即时有效检验与置信序列。利用赌局/e-过程框架,将时域感知的下注策略建模为以 $(t, \log W_t)$ 为状态空间的有限时域最优控制问题,其中 $t$ 为时间,$W_t$ 为检验鞅值。我们首先证明,在状态空间的某些内部区域,显著偏离凯利下注的策略在理论上必劣于凯利下注,后者能以高概率达到阈值。随后识别出充分条件:当处于该区域之外时,若下注者进度落后,更激进的下注可能更优;若进度领先,则更保守的下注更佳。这些结果共同构成 $(t, \log W_t)$ 平面上的简单分阶段图示,划分出凯利、部分凯利及激进下注的适用区域。基于此分阶段图示,我们提出一种基于通用深度 Q 网络(DQN)的强化学习方法,该智能体从合成经验中学习单一策略,并将过去观测的简单统计量映射为跨时域与零假设值的下注策略。在有限时域实验中,所学的 DQN 策略取得了当前最佳性能。
原文摘要 · Abstract (English)
We develop horizon-aware anytime-valid tests and confidence sequences for bounded means under a strict deadline $N$. Using the betting/e-process framework, we cast horizon-aware betting as a finite-horizon optimal control problem with state space $(t, \log W_t)$, where $t$ is the time and $W_t$ is the test martingale value. We first show that in certain interior regions of the state space, policies that deviate significantly from Kelly betting are provably suboptimal, while Kelly betting reaches the threshold with high probability. We then identify sufficient conditions showing that outside this region, more aggressive betting than Kelly can be better if the bettor is behind schedule, and less aggressive can be better if the bettor is ahead. Taken together these results suggest a simple phase diagram in the $(t, \log W_t)$ plane, delineating regions where Kelly, fractional Kelly, and aggressive betting may be preferable. Guided by this phase diagram, we introduce a Deep Reinforcement Learning approach based on a universal Deep Q-Network (DQN) agent that learns a single policy from synthetic experience and maps simple statistics of past observations to bets across horizons and null values. In limited-horizon experiments, the learned DQN policy yields state-of-the-art results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。