用概率方法让智能体在想象中高效试错,提升世界模型学习效果。
Probabilistic Dreaming for World Models
- 通过并行探索多个潜在状态,增强想象力的多样性。
- 在简单标签任务中得分提高4.5%,收益方差降低28%。
- 适合研究高效强化学习与不确定性建模的学者。
想象式学习使智能体能从虚构经验中汲取知识,从而更稳健、高效地学习世界模型。本文对当前领先的Dreamer模型引入概率方法,实现:(1) 并行探索多个潜在状态;(2) 在保持连续潜在变量梯度优势的同时,为互斥未来保留不同假设。在MPE SimpleTag环境中评估,该方法相比标准Dreamer获得4.5%的分数提升,且每回合回报方差降低28%。我们还讨论了局限性与未来方向,包括最优超参数(如粒子数K)如何随环境复杂度变化,以及在世界模型中捕捉认知不确定性的方法。
原文摘要 · Abstract (English)
"Dreaming" enables agents to learn from imagined experiences, enabling more robust and sample-efficient learning of world models. In this work, we consider innovations to the state-of-the-art Dreamer model using probabilistic methods that enable: (1) the parallel exploration of many latent states; and (2) maintaining distinct hypotheses for mutually exclusive futures while retaining the desirable gradient properties of continuous latents. Evaluating on the MPE SimpleTag domain, our method outperforms standard Dreamer with a 4.5% score improvement and 28% lower variance in episode returns. We also discuss limitations and directions for future work, including how optimal hyperparameters (e.g. particle count K) scale with environmental complexity, and methods to capture epistemic uncertainty in world models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。