用线性时间建模提升扩散世界模型的内存能力,显著增强强化学习表现。
EDELINE: Enhancing Memory in Diffusion-based World Models via Linear-Time Sequence Modeling
- 融合状态空间模型与扩散模型,实现线性时间序列建模
- 在Atari 100k、Crafter和ViZDoom中均超越现有基线
- 适合需要长程记忆的复杂视觉强化学习任务
世界模型为强化学习代理提供了极高的样本效率。现有方法多依赖离散潜变量序列建模环境动态,但这种压缩常忽略强化学习所需的关键视觉细节。近期基于扩散的世界模型通过固定长度帧上下文预测下一观察,使用独立的循环神经网络建模奖励与终止信号。尽管该架构提升了视觉保真度,但固定上下文长度限制了记忆容量。本文提出EDELINE,一种将状态空间模型与扩散模型统一的架构。其在视觉挑战性强的Atari 100k任务、对记忆要求高的Crafter基准以及3D第一人称ViZDoom环境中均优于现有基线,展现出全面优越性能。
原文摘要 · Abstract (English)
World models represent a promising approach for training reinforcement learning agents with significantly improved sample efficiency. While most world model methods primarily rely on sequences of discrete latent variables to model environment dynamics, this compression often neglects critical visual details essential for reinforcement learning. Recent diffusion-based world models condition generation on a fixed context length of frames to predict the next observation, using separate recurrent neural networks to model rewards and termination signals. Although this architecture effectively enhances visual fidelity, the fixed context length approach inherently limits memory capacity. In this paper, we introduce EDELINE, a unified world model architecture that integrates state space models with diffusion models. Our approach outperforms existing baselines across visually challenging Atari 100k tasks, memory-demanding Crafter benchmark, and 3D first-person ViZDoom environments, demonstrating superior performance in all these diverse challenges.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。