针对高随机环境的探索,提出按方差分配采样资源的新方法。
Variance Driven Exploration: A Provable and Efficient Methodology for Pure Exploration in Highly Stochastic Environments

- 根据各组件噪声对最终决策的影响,动态分配采样资源以降低不确定性
- 在三类纯探索任务中实现方差衰减与简单后悔率的理论保证
- 在高随机环境中显著优于现有方法,适合强化学习与贝叶斯优化场景
我们提出方差驱动探索(VarDE),一种在高度随机环境中进行纯探索的原理性方法。该方法的核心思想是:采样资源应被分配以最小化最终决策的不确定性。通过引入平滑的决策函数来形式化这种不确定性,并推导出明确捕捉各组件随机噪声如何影响最终输出可靠性的资源分配规则。我们将该方法应用于纯探索的三个核心问题——最佳臂识别(BAI)、蒙特卡洛树搜索(MCTS)和最佳策略识别(BPI),并提供了关于方差衰减和简单后悔率的理论保证。实验表明,相较于现有方法,VarDE在高随机环境下表现出持续且显著的性能提升。
原文摘要 · Abstract (English)
We propose Variance Driven Exploration (VarDE), a principled approach for pure exploration in highly stochastic environments, where the exploration process is dominated by stochastic variance. VarDE is built on a fundamental principle: sampling effort should be allocated to minimize the uncertainty of the final decision. We formalize the uncertainty of the final decision through a smooth decision function and derive allocation rules that explicitly capture how stochastic noise in individual components affects the reliability of the final output. We apply this methodology to three core problems of pure exploration -- Best Arm Identification (BAI), Monte Carlo Tree Search (MCTS), and Best-Policy Identification (BPI) -- with theoretical guarantees on variance decay and simple regret. Empirically, we demonstrate consistent and significant improvements of VarDE over existing methods, with especially strong gains in highly stochastic environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。