不依赖RNN、Transformer和图像重建,用简单高效方法构建世界模型。
Simple, Good, Fast: Self-Supervised World Models Free of Baggage
- 用自监督学习和帧/动作堆叠捕捉短期依赖
- 在Atari 100k上表现优异,优于多个基线模型
- 适合追求轻量高效世界模型的研究者
世界模型的核心构成是什么?若摒弃RNN、Transformer、离散表示和图像重建,仍能取得多好效果?本文提出SGF——一种简单、高效且性能优良的世界模型。它采用自监督表示学习,通过帧与动作堆叠捕捉短时依赖,并利用数据增强提升对模型误差的鲁棒性。我们深入讨论了SGF与已有世界模型的联系,通过消融实验评估各模块作用,并在Atari 100k基准上通过定量对比验证其良好表现。
原文摘要 · Abstract (English)
What are the essential components of world models? How far do we get with world models that are not employing RNNs, transformers, discrete representations, and image reconstructions? This paper introduces SGF, a Simple, Good, and Fast world model that uses self-supervised representation learning, captures short-time dependencies through frame and action stacking, and enhances robustness against model errors through data augmentation. We extensively discuss SGF's connections to established world models, evaluate the building blocks in ablation studies, and demonstrate good performance through quantitative comparisons on the Atari 100k benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。