在未知系统中,优化并行与重启策略可显著提升稀有状态探索效率。
Efficiency of Parallel and Restart Exploration Strategies in Model Free Stochastic Simulations
- 通过随机游走和莱维过程建模探索行为,分析并行数量对成功率的影响。
- 存在最优并行数 $N^*$,超过后成功率指数下降,证明资源分配需权衡。
- 重启机制可将停滞轨迹资源重分配,使成功概率实现指数级提升,适合强化学习应用。
我们分析了在模型无关设置下,针对未知系统动态的随机模拟中,并行化与重启机制的效率问题。这类场景常见于强化学习和罕见事件估计,标准方差缩减技术(如重要性采样)无法适用。聚焦于有限计算预算下到达稀有状态的挑战,我们采用随机游走和莱维过程建模探索行为。基于严格的概率分析,揭示了成功概率随并行模拟数量变化的相变现象:存在一个最优并行数 $N^*$,在该点平衡探索多样性与每条轨迹的时间分配;超过此阈值后,性能呈指数级下降。此外,我们证明了一种重启策略——将停滞轨迹的资源重新分配至有前景区域——可带来成功概率的指数级提升。在强化学习中,这些策略能改进策略梯度方法,实现更高效的状空间探索,从而获得更准确的策略梯度估计。
原文摘要 · Abstract (English)
We analyze the efficiency of parallelization and restart mechanisms for stochastic simulations in model-free settings, where the underlying system dynamics are unknown. Such settings are common in Reinforcement Learning (RL) and rare event estimation, where standard variance-reduction techniques like importance sampling are inapplicable. Focusing on the challenge of reaching rare states under a finite computational budget, we model exploration via random walks and Lévy processes. Based on rigorous probability analysis, our work reveals a phase transition in the success probability as a function of the number of parallel simulations: an optimal number $N^*$ exists, balancing exploration diversity and time allocation per simulation. Beyond this threshold, performance degrades exponentially. Furthermore, we demonstrate that a restart strategy, which reallocates resources from stagnant trajectories to promising regions, can yield an exponential improvement in success probability. In the context of RL, these strategies can improve policy gradient methods by enabling more efficient state-space exploration, leading to more accurate policy gradient estimates.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。