让模型跳出历史数据,主动探索未知领域,加速稀疏奖励任务学习。
Mind Dreamer: Untethering Imagination via Active Causal Intervention on Latent Manifolds

- 用对抗生成器初始化潜空间状态,实现非连续的主动干预。
- 在稀疏奖励任务中速度提升8.8倍,平均快1.67倍。
- 适合需要高效探索的强化学习场景,尤其稀疏奖励环境。
基于模型的强化学习通过潜空间想象实现采样高效,但受限于历史依赖:想象通常从观测状态开始。这导致世界模型的流形发现快于策略的稀疏奖励优化,形成学习不对称。本文提出 Mind Dreamer (MD),通过主动因果干预突破马尔可夫连续性限制。MD 将发现过程重构为全局中继期望自由能最小化。不依赖历史数据初始化,而是从对抗生成器 $s_0 ilde{p}_{gen}(ullet)$ 中采样初始状态,实现物理合理但认知挑战性的潜空间跳跃。引入中继价值函数与中继不确定性函数,解决跨空间断裂的信用分配难题。将合成锚点视为干预中介状态,通过贝尔曼风格回溯传播实用与认知价值。理论证明,不确定性在断点间传播需二次折扣 $γ^2$,确立形式化的认知边界。理论上,MD 近似于方差最小化重要性采样器,扩大流形谱间隙,降低到达关键瓶颈状态的命中时间。实验上,MD 在 DeepMind Control Suite 上平均提速 1.67×,稀疏奖励任务达 8.8×。
原文摘要 · Abstract (English)
Model-Based Reinforcement Learning yields sample efficiency via latent imagination, yet remains constrained by Historical Tethering: imagination is typically initialized from observed states. This creates a learning asymmetry, where the world model's manifold discovery outpaces the policy's sparse-reward optimization. We propose Mind Dreamer (MD), a framework that instantiates Active Causal Intervention to transcend Markovian continuity. MD reformulates discovery as the minimization of a global Relay Expected Free Energy. Instead of initializing from historical data, it draws initial states from an adversarial generator $s_0 \sim p_{gen}(\cdot)$, creating non-continuous latent jumps to epistemic blind spots that are physically plausible yet cognitively challenging. We derive Relay Value Function and Relay Uncertainty Function to resolve the credit assignment paradox across these spatial ruptures. Treating synthesized anchors as interventional intermediary states, these potentials propagate pragmatic and epistemic value through Bellman-style backups. Notably, we prove that uncertainty propagation across discontinuities necessitates a quadratic discount $γ^2$, establishing a formal epistemic horizon. Theoretically, MD approximates a variance-minimizing importance sampler that expands the manifold's spectral gap, reducing the hitting time to critical bottleneck states. Empirically, MD achieves a 1.67$\times$ average speedup over DreamerV3 on DeepMind Control Suite, reaching 8.8$\times$ in sparse-reward tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。