无需重建图像,用连续确定性表示预测提升世界模型性能
Dreamer-CDP: Improving Reconstruction-free World Models Via Continuous Deterministic Representation Prediction
- 设计连续确定性表示预测器替代图像重建
- 在Crafter环境达到与重建方法相当的性能
- 适合追求高效、无重建依赖的世界模型研究者
基于模型的强化学习(MBRL)在高维观测空间中,如Dreamer,依赖抽象表征进行有效规划与控制。现有方法通常采用观测空间的重建目标,导致表征对任务无关细节敏感。近期替代方案虽放弃重建,改用辅助动作预测头或视图增强策略,但在Crafter环境中的表现仍不如重建方法。本文提出一种类似JEPA的预测器,作用于连续、确定性的表征上,成功弥补了这一差距。所提方法在Crafter基准上达到与Dreamer相当的性能,证明了无需重建目标也能实现有效的世界模型学习。
原文摘要 · Abstract (English)
Model-based reinforcement learning (MBRL) agents operating in high-dimensional observation spaces, such as Dreamer, rely on learning abstract representations for effective planning and control. Existing approaches typically employ reconstruction-based objectives in the observation space, which can render representations sensitive to task-irrelevant details. Recent alternatives trade reconstruction for auxiliary action prediction heads or view augmentation strategies, but perform worse in the Crafter environment than reconstruction-based methods. We close this gap between Dreamer and reconstruction-free models by introducing a JEPA-style predictor defined on continuous, deterministic representations. Our method matches Dreamer's performance on Crafter, demonstrating effective world model learning on this benchmark without reconstruction objectives.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。