用时序变换器预测嵌入,让世界模型更懂时间逻辑。
Next Embedding Prediction Makes World Models Stronger
- 用时序变换器直接预测下一步编码嵌入,无需重建损失
- 在深脑控制套件上性能媲美甚至超越DreamerV3
- 特别适合需要记忆与空间推理的复杂任务
在部分可观测、高维环境中,捕捉时间依赖性对基于模型的强化学习(MBRL)至关重要。我们提出NE-Dreamer,一种无解码器的MBRL智能体,利用时序变换器从潜在状态序列中预测下一步编码嵌入,直接优化表示空间中的时间预测一致性。该方法使NE-Dreamer无需重构损失或辅助监督即可学习连贯且具有预测性的状态表示。在DeepMind Control Suite上,NE-Dreamer性能匹配或超越DreamerV3及领先无解码器智能体。在包含记忆与空间推理挑战的DMLab子集任务中,性能显著提升。这些结果证明,结合时序变换器的下一嵌入预测是复杂部分可观测环境中有效且可扩展的MBRL框架。
原文摘要 · Abstract (English)
Capturing temporal dependencies is critical for model-based reinforcement learning (MBRL) in partially observable, high-dimensional domains. We introduce NE-Dreamer, a decoder-free MBRL agent that leverages a temporal transformer to predict next-step encoder embeddings from latent state sequences, directly optimizing temporal predictive alignment in representation space. This approach enables NE-Dreamer to learn coherent, predictive state representations without reconstruction losses or auxiliary supervision. On the DeepMind Control Suite, NE-Dreamer matches or exceeds the performance of DreamerV3 and leading decoder-free agents. On a challenging subset of DMLab tasks involving memory and spatial reasoning, NE-Dreamer achieves substantial gains. These results establish next-embedding prediction with temporal transformers as an effective, scalable framework for MBRL in complex, partially observable environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。