让智能体在未知环境下零样本泛化,靠隐式学习环境动态
Dynamics-Aligned Latent Imagination in Contextual World Models for Zero-Shot Generalization
- 通过自监督编码器从交互中推断隐藏环境状态
- 在多个挑战性任务上超越依赖显式环境变量的方法
- 适合需要快速适应新环境的强化学习场景
现实世界的强化学习要求智能体在不重新训练的情况下适应未见过的环境条件。上下文马尔可夫决策过程(cMDP)描述了这一挑战,但现有方法通常依赖显式的上下文变量(如摩擦力、重力),当这些变量为隐式或难以测量时便受限。我们提出动态对齐隐式想象(DALI),集成于Dreamer架构中,通过自监督编码器从智能体-环境交互中推断隐式上下文表征。该编码器训练目标是预测前向动态,生成可作用于世界模型与策略的上下文表示,实现感知与控制的融合。我们理论上证明该编码器对高效上下文推断和鲁棒泛化至关重要。DALI的隐空间具备反事实一致性:扰动重力编码维度会以物理合理的方式改变想象轨迹。在多个挑战性的cMDP基准测试中,DALI显著优于无上下文感知基线,常在外推任务中超越依赖显式上下文的基线,实现对未见上下文变化的零样本泛化。
原文摘要 · Abstract (English)
Real-world reinforcement learning demands adaptation to unseen environmental conditions without costly retraining. Contextual Markov Decision Processes (cMDP) model this challenge, but existing methods often require explicit context variables (e.g., friction, gravity), limiting their use when contexts are latent or hard to measure. We introduce Dynamics-Aligned Latent Imagination (DALI), a framework integrated within the Dreamer architecture that infers latent context representations from agent-environment interactions. By training a self-supervised encoder to predict forward dynamics, DALI generates actionable representations conditioning the world model and policy, bridging perception and control. We theoretically prove this encoder is essential for efficient context inference and robust generalization. DALI's latent space enables counterfactual consistency: Perturbing a gravity-encoding dimension alters imagined rollouts in physically plausible ways. On challenging cMDP benchmarks, DALI achieves significant gains over context-unaware baselines, often surpassing context-aware baselines in extrapolation tasks, enabling zero-shot generalization to unseen contextual variations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。