通过谱约束动态建模,提升长时幻想轨迹的稳定性与控制性能
Koopman Dreamer: Spectrally Constrained Latent Dynamics for Stable World-Model Imagination

- 采用二维旋转-缩放块构建谱约束的确定性潜空间动力学
- 在连续控制任务中实现更稳定的长程幻想,误差累积显著降低
- 适合需要高精度多步预测的机器人控制场景
潜空间世界模型通过在想象的潜轨迹上优化策略,提升了连续控制中的样本效率,但常见的神经动力学难以直接控制模态持续性和长期回溯中的误差积累。我们提出Koopman Dreamer,一种基于Dreamer架构、带有谱约束的确定性潜空间动力学核心的世界模型。其受Koopman理论启发的主干结构采用二维旋转-缩放块,以有界半径表示阻尼、旋转和近周期模式;线性和低秩双线性动作项捕捉全局与状态依赖的控制效应,随机状态调制提供局部校正信息。为缓解后验条件训练与先验仅幻想之间的偏差,模型结合后验条件的EMA教师目标与一步一致性、多步回溯及开环观测预测目标。进一步推导出多步回溯误差上界,将谱主干引起的放大效应与双线性交互分离,并与随机状态不匹配和建模残差的加性效应区分开来,揭示了误差衰减与长期信息保留间的权衡。在DeepMind Control Suite的本体感觉连续控制任务及无人机激光雷达自主导航任务上的实验表明,Koopman Dreamer显著提升了长时潜回溯的稳定性,在依赖高质量多步幻想的任务中实现了更强的闭环控制性能。
原文摘要 · Abstract (English)
Latent world models improve sample efficiency in continuous control by optimizing policies over imagined latent trajectories, but common neural transitions offer limited direct control over modal persistence and error accumulation in long rollouts. We propose Koopman Dreamer, a Dreamer-style world model with a spectrally constrained deterministic latent dynamics core. Its Koopman-inspired backbone uses two-dimensional rotation--scaling blocks with bounded radii to represent damping, rotation, and near-periodic modes. Linear and low-rank bilinear action terms capture global and state-dependent control effects, while stochastic-state modulation supplies local correction information. To reduce the mismatch between posterior-conditioned training and prior-only imagination, the model combines posterior-conditioned EMA teacher targets with one-step consistency, multi-step rollout, and open-loop observation-prediction objectives. We further derive a multi-step rollout-error bound that separates amplification by the spectral backbone and bilinear interaction from the additive effects of stochastic-state mismatch and modeling residuals, clarifying the trade-off between error attenuation and long-term information retention. Experimental results on proprioceptive continuous-control tasks from the DeepMind Control Suite and UAV-LiDAR autonomous navigation demonstrate that Koopman Dreamer improves the stability of long-horizon latent rollouts and achieves stronger closed-loop control performance on tasks that rely on high-quality multi-step imagination.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。