arXiv:2509.19305cs.LGcs.AI2025-09被引 3

通过频域分析提升强化学习轨迹稳定性

Wavelet Fourier Diffuser: Frequency-Aware Diffusion Model for Reinforcement Learning

  • 用小波变换分离轨迹的高低频成分
  • 结合短时傅里叶变换与交叉注意力提取频域特征
  • 在D4RL上实现更平滑稳定、性能更优的轨迹

扩散概率模型在离线强化学习中展现出显著潜力,可直接建模轨迹序列。然而,现有方法主要关注时域特征,忽略频域特征,导致频率偏移并降低性能。本文从频域角度重新审视强化学习问题,发现仅使用时域的方法会无意引入频域低频分量的偏移,造成轨迹不稳定和性能下降。为此,我们提出一种新型基于扩散的强化学习框架——小波傅里叶扩散器(WFDiffuser),该框架利用离散小波变换将轨迹分解为低频与高频成分,并采用短时傅里叶变换与交叉注意力机制提取频域特征,促进跨频段交互。在D4RL基准上的大量实验表明,WFDiffuser有效缓解了频率偏移,生成的轨迹更平滑、更稳定,决策性能优于现有方法。

原文摘要 · Abstract (English)

Diffusion probability models have shown significant promise in offline reinforcement learning by directly modeling trajectory sequences. However, existing approaches primarily focus on time-domain features while overlooking frequency-domain features, leading to frequency shift and degraded performance according to our observation. In this paper, we investigate the RL problem from a new perspective of the frequency domain. We first observe that time-domain-only approaches inadvertently introduce shifts in the low-frequency components of the frequency domain, which results in trajectory instability and degraded performance. To address this issue, we propose Wavelet Fourier Diffuser (WFDiffuser), a novel diffusion-based RL framework that integrates Discrete Wavelet Transform to decompose trajectories into low- and high-frequency components. To further enhance diffusion modeling for each component, WFDiffuser employs Short-Time Fourier Transform and cross attention mechanisms to extract frequency-domain features and facilitate cross-frequency interaction. Extensive experiment results on the D4RL benchmark demonstrate that WFDiffuser effectively mitigates frequency shift, leading to smoother, more stable trajectories and improved decision-making performance over existing methods.

强化学习扩散模型频域分析轨迹生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。