arXiv:2601.19624cs.LGcs.AI2026-01中稿 · ICML被引 5

根据环境变化程度动态调整探索强度,提升强化学习在变化环境中的适应性。

Tracking Drift: Variation-Aware Entropy Scheduling for Non-Stationary Reinforcement Learning

  • 基于在线漂移指标自适应调节熵系数,实现动态探索
  • 在12个任务中显著减少漂移导致的性能下降,加速恢复
  • 无需修改算法结构,计算开销极小,适合实际部署

现实世界中的强化学习常面临环境漂移问题,但现有方法多依赖固定的熵系数或目标熵,导致稳定期过度探索、漂移后探索不足,且未回答探索强度如何随漂移幅度变化的理论问题。我们证明,在标准假设下,非平稳最大熵强化学习中的熵调度可转化为跟踪漂移参考与更新稳定性的动态后悔权衡,从而导出熵权重与在线非平稳性代理的平方根缩放关系。基于此,我们提出AES——自适应熵调度方法,通过训练过程中可观测的漂移代理在线调整熵系数/温度,几乎无需结构改动且开销极小。在4种算法变体、12个任务和4种漂移模式下,AES显著降低了漂移引起的性能下降比例,并加快了突变后的恢复速度。

原文摘要 · Abstract (English)

Real-world reinforcement learning often faces environment drift, but most existing methods rely on static entropy coefficients/target entropy, causing over-exploration during stable periods and under-exploration after drift, and leaving unanswered the principled question of how exploration intensity should scale with drift magnitude. We show that, under standard assumptions, entropy scheduling in non-stationary maximum-entropy RL can be cast as the dynamic-regret trade-off between tracking a drifting comparator and stabilizing updates, yielding a square-root scaling rule for the entropy weight in terms of a online non-stationarity proxy. Building on this, we propose AES--Adaptive Entropy Scheduling--which adaptively adjusts the entropy coefficient/temperature online using observable drift proxies during training, requiring almost no structural changes and incurring minimal overhead. Across 4 algorithm variants, 12 tasks, and 4 drift modes, AES significantly reduces the fraction of performance degradation caused by drift and accelerates recovery after abrupt changes.

强化学习自适应调度非平稳性探索效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。