arXiv:2505.11985cs.LGstat.ML2025-05被引 1

找方差最大的动作,优化选择策略以减少错误分配和提升识别精度。

Variance-Optimal Arm Selection: Misallocation Minimization and Best Arm Identification

  • 提出UCB-VV算法,通过上界控制错误选择次数随时间对数增长。
  • 在固定预算下SHVV算法误差率呈指数下降,理论性能达到最优。
  • 首次推导子高斯分布下样本方差与夏普比率的集中不等式,适用于金融交易场景。

本文研究从K个独立动作中选出方差最高的动作。针对两种情形:(i) 错误分配最小化(MM),惩罚次优动作的选取次数;(ii) 固定预算最佳动作识别(BAI),评估算法在固定抽取次数后正确识别最高方差动作的能力。我们提出新型在线算法UCB-VV用于MM设置,证明其错误分配上界为$/mathcal{O}(/log{n})$,其中$n$为时间范围,并通过下界证明该算法阶次最优。对于固定预算下的BAI,提出SHVV算法,其误差概率上界为$/expig(-n/(/log(K) H)ig)$,其中$H$表示问题复杂度,该速率与对应下界一致。将框架扩展至子高斯分布,利用新推导的样本方差与标准差集中不等式,首次获得子高斯分布下经验夏普比率的集中不等式。实验显示UCB-VV在不同次优差距下均优于ε-贪心,虽略逊于无理论保证的VTS;SHVV在6种设置下优于均匀采样。案例研究在100只基于GBM生成的股票期权交易中验证了两算法的有效性。

原文摘要 · Abstract (English)

This paper focuses on selecting the arm with the highest variance from a set of $K$ independent arms. Specifically, we focus on two settings: (i) misallocation minimization setting, that penalizes the number of pulls of suboptimal arms in terms of variance, and (ii) fixed-budget best arm identification setting, that evaluates the ability of an algorithm to determine the arm with the highest variance after a fixed number of pulls. We develop a novel online algorithm called UCB-VV for the misallocation minimization (MM) and show that its upper bound on misallocation for bounded rewards evolves as $\mathcal{O}\left(\log{n}\right)$ where $n$ is the horizon. By deriving the lower bound on the misallocation, we show that UCB-VV is order optimal. For the fixed budget best arm identification (BAI) setting we propose the SHVV algorithm. We show that the upper bound of the error probability of SHVV evolves as $\exp\left(-\frac{n}{\log(K) H}\right)$, where $H$ represents the complexity of the problem, and this rate matches the corresponding lower bound. We extend the framework from bounded distributions to sub-Gaussian distributions using a novel concentration inequality on the sample variance and standard deviation. Leveraging the same, we derive a concentration inequality for the empirical Sharpe ratio (SR) for sub-Gaussian distributions, which was previously unknown in the literature. Empirical simulations show that UCB-VV consistently outperforms $ε$-greedy across different sub-optimality gaps though it is surpassed by VTS, which exhibits the lowest misallocation, albeit lacking in theoretical guarantees. We also illustrate the superior performance of SHVV, for a fixed budget setting under 6 different setups against uniform sampling. Finally, we conduct a case study to empirically evaluate the performance of the UCB-VV and SHVV in call option trading on $100$ stocks generated using GBM.

强化学习最优选择金融应用概率不等式

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。