用世界模型生成参考轨迹,实现快速运动适应。
World Models as Reference Trajectories for Rapid Motor Adaptation
- 双控架构:强化学习优化长期奖励,潜空间快速执行。
- 适应速度远超基线,计算开销低且性能接近最优。
- 适合动态变化的高维连续控制任务,如机器人操控。
将学习到的控制策略部署到真实环境面临根本挑战:当系统动力学意外改变时,性能会下降,直到在新数据上重新训练模型。我们提出反射式世界模型(RWM),一种双控框架,利用世界模型预测作为隐式参考轨迹,实现快速适应。该方法将控制问题分解为通过强化学习实现的长期奖励最大化,以及通过快速潜空间控制实现的鲁棒运动执行。相比基于模型的强化学习基线,该双架构在保持近最优性能的同时,实现了显著更快的适应速度,且在线计算成本极低。该方法结合了强化学习的灵活策略学习与快速误差校正能力,为在不同动力学条件下维持高性能提供了一种原则性解决方案,适用于高维连续控制任务。
原文摘要 · Abstract (English)
Deploying learned control policies in real-world environments poses a fundamental challenge. When system dynamics change unexpectedly, performance degrades until models are retrained on new data. We introduce Reflexive World Models (RWM), a dual control framework that uses world model predictions as implicit reference trajectories for rapid adaptation. Our method separates the control problem into long-term reward maximization through reinforcement learning and robust motor execution through rapid latent control. This dual architecture achieves significantly faster adaptation with low online computational cost compared to model-based RL baselines, while maintaining near-optimal performance. The approach combines the benefits of flexible policy learning through reinforcement learning with rapid error correction capabilities, providing a principled approach to maintaining performance in high-dimensional continuous control tasks under varying dynamics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。