MAGE通过多尺度生成提升离线强化学习长时稀疏奖励任务表现
MAGE: Multi-scale Autoregressive Generation for Offline Reinforcement Learning
- 构建多尺度自回归生成框架,分层建模轨迹时间结构
- 在5个基准上超越15种基线,长时轨迹生成更连贯可控
- 适合需要精细短期行为控制的复杂决策场景
生成模型因能刻画复杂轨迹分布而在离线强化学习中备受关注。然而,现有基于生成的方法在长时程、稀疏奖励任务上仍表现不佳。一些分层生成方法通过单一策略分解为短时子问题,并用另一策略生成细节动作来缓解此问题,但常忽略轨迹固有的多尺度时间结构,导致性能受限。为此,我们提出MAGE——一种基于多尺度自回归生成的离线强化学习方法。MAGE采用条件引导的多尺度自编码器学习层次化轨迹表征,并利用多尺度Transformer从粗到细自回归生成轨迹表征,有效捕捉多分辨率下的时间依赖性。同时,引入条件引导解码器以精确控制短期行为。在五个离线强化学习基准上与十五种基线算法对比的大量实验表明,MAGE成功将多尺度轨迹建模与条件引导相结合,在长时程稀疏奖励设定下生成了连贯且可控制的轨迹。
原文摘要 · Abstract (English)
Generative models have gained significant traction in offline reinforcement learning (RL) due to their ability to model complex trajectory distributions. However, existing generation-based approaches still struggle with long-horizon tasks characterized by sparse rewards. Some hierarchical generation methods have been developed to mitigate this issue by decomposing the original problem into shorter-horizon subproblems using one policy and generating detailed actions with another. While effective, these methods often overlook the multi-scale temporal structure inherent in trajectories, resulting in suboptimal performance. To overcome these limitations, we propose MAGE, a Multi-scale Autoregressive GEneration-based offline RL method. MAGE incorporates a condition-guided multi-scale autoencoder to learn hierarchical trajectory representations, along with a multi-scale transformer that autoregressively generates trajectory representations from coarse to fine temporal scales. MAGE effectively captures temporal dependencies of trajectories at multiple resolutions. Additionally, a condition-guided decoder is employed to exert precise control over short-term behaviors. Extensive experiments on five offline RL benchmarks against fifteen baseline algorithms show that MAGE successfully integrates multi-scale trajectory modeling with conditional guidance, generating coherent and controllable trajectories in long-horizon sparse-reward settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。