arXiv:2603.16842cs.LGcond-mat.dis-nn2026-03

随机重置可加速强化学习,通过重塑经验积累方式提升学习效率。

Stochastic Resetting Accelerates Reinforcement Learning Beyond Random Search

  • 引入随机重置机制,让智能体周期性回到初始状态以加速学习。
  • 在网格任务中,重置使奖励信息传播更快,即使不缩短搜索时间。
  • 适合探索困难、奖励稀疏的复杂任务,且不改变最终最优策略。

随机重置——间歇性将过程返回到固定参考状态——已成为优化首次通过特性的一种有效机制。现有理论主要针对不学习的搜索过程:搜索者遵循固定动力学,在重置之间不积累知识。本文探讨随机重置如何与强化学习(其中底层动力学通过经验自适应)相互作用。在表格型网格环境中,我们发现即使重置未减少扩散智能体的搜索时间,也能加速学习。结果揭示了一种新的机制:重置通过加速奖励信息的传播来促进学习。我们发现,确定性、尖锐的重置比随机协议在更窄的重置率范围内加速学习。在使用神经网络价值近似的连续状态任务中,我们证明当探索困难且奖励稀疏时,重置能显著加快学习速度。进一步指出,在表格任务中,重置不改变最终解,不同于时间折扣等会偏移最优行为的技术。结果确立了随机重置作为一种简单、可调的机制,通过塑造经验累积方式加速学习,将统计力学中的经典现象拓展至自适应系统。

原文摘要 · Abstract (English)

Stochastic resetting -- intermittently returning a process to a fixed reference state -- has emerged as an effective mechanism for optimizing first-passage properties. Existing theory largely treats processes that search but do not learn: the searcher follows fixed dynamics, accumulating no knowledge between resets. Here we ask how stochastic resetting interacts with reinforcement learning, where the underlying dynamics adapt through experience. In tabular grid environments, we find that resetting can accelerate learning even when it does not reduce the search time of a diffusive agent. Our results reveal a distinct additional mechanism through which resetting speeds the propagation of reward information. We show that deterministic, sharp resetting accelerates learning more than the stochastic protocol but over a narrower range of reset rates. In a continuous-state task with neural-network-based value approximation, we demonstrate that resetting speeds up learning when exploration is hard and rewards are sparse. We argue further that, in the tabular tasks, resetting accelerates learning without altering the solution the agent ultimately reaches, unlike other techniques such as temporal discounting, which biases the optimal behavior. Our results establish stochastic resetting as a simple, tunable mechanism for accelerating learning by shaping how experience accumulates, extending a canonical phenomenon of statistical mechanics to adaptive systems.

强化学习随机重置学习加速动态系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。