arXiv:2510.04507cs.LG2025-10

用小波分析提升强化学习在动态环境中的适应能力

Wavelet Predictive Representations for Non-Stationary Reinforcement Learning

  • 将任务表示转至小波域,捕捉多尺度变化特征
  • 小波TD更新使策略在非平稳环境中更快收敛
  • 适合复杂动态场景下的快速适应型强化学习

现实世界具有内在非平稳性,如天气和交通流的持续变化,给智能体适应环境动态带来挑战。非平稳强化学习(NSRL)旨在训练智能体快速适应一系列不同的马尔可夫决策过程(MDPs)。然而,现有方法多聚焦于规律变化的任务,难以应对高度动态环境。受小波分析在时间序列建模中成功启发,本文提出WISDOM,利用小波域的预测性表征增强NSRL。WISDOM通过将任务表示序列变换到小波域,使小波系数同时表征全局趋势与细粒度变化。除常规自回归建模外,设计了小波时序差分(TD)更新算子,以提升对MDP演化的跟踪与预测能力。理论证明该算子收敛,并实现策略改进。在多个基准测试中,WISDOM显著优于现有基线,在样本效率和最终性能上均表现优异,展现了在非平稳、随机演化任务中的强大适应能力。

原文摘要 · Abstract (English)

The real world is inherently non-stationary, with ever-changing factors, such as weather conditions and traffic flows, making it challenging for agents to adapt to varying environmental dynamics. Non-Stationary Reinforcement Learning (NSRL) addresses this challenge by training agents to adapt rapidly to sequences of distinct Markov Decision Processes (MDPs). However, existing NSRL approaches often focus on tasks with regularly evolving patterns, leading to limited adaptability in highly dynamic settings. Inspired by the success of Wavelet analysis in time series modeling, specifically its ability to capture signal trends at multiple scales, we propose WISDOM to leverage wavelet-domain predictive task representations to enhance NSRL. WISDOM captures these multi-scale features in evolving MDP sequences by transforming task representation sequences into the wavelet domain, where wavelet coefficients represent both global trends and fine-grained variations of non-stationary changes. In addition to the auto-regressive modeling commonly employed in time series forecasting, we devise a wavelet temporal difference (TD) update operator to enhance tracking and prediction of MDP evolution. We theoretically prove the convergence of this operator and demonstrate policy improvement with wavelet task representations. Experiments on diverse benchmarks show that WISDOM significantly outperforms existing baselines in both sample efficiency and asymptotic performance, demonstrating its remarkable adaptability in complex environments characterized by non-stationary and stochastically evolving tasks.

强化学习非平稳小波分析自适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。