arXiv:2505.15589cs.LGcs.AI2025-05NeurIPS被引 2

用世界模型生成参考轨迹,实现快速运动适应。

World Models as Reference Trajectories for Rapid Motor Adaptation

  • 双控架构:强化学习优化长期奖励,潜空间快速执行。
  • 适应速度远超基线,计算开销低且性能接近最优。
  • 适合动态变化的高维连续控制任务,如机器人操控。

将学习到的控制策略部署到真实环境面临根本挑战:当系统动力学意外改变时,性能会下降,直到在新数据上重新训练模型。我们提出反射式世界模型(RWM),一种双控框架,利用世界模型预测作为隐式参考轨迹,实现快速适应。该方法将控制问题分解为通过强化学习实现的长期奖励最大化,以及通过快速潜空间控制实现的鲁棒运动执行。相比基于模型的强化学习基线,该双架构在保持近最优性能的同时,实现了显著更快的适应速度,且在线计算成本极低。该方法结合了强化学习的灵活策略学习与快速误差校正能力,为在不同动力学条件下维持高性能提供了一种原则性解决方案,适用于高维连续控制任务。

原文摘要 · Abstract (English)

Deploying learned control policies in real-world environments poses a fundamental challenge. When system dynamics change unexpectedly, performance degrades until models are retrained on new data. We introduce Reflexive World Models (RWM), a dual control framework that uses world model predictions as implicit reference trajectories for rapid adaptation. Our method separates the control problem into long-term reward maximization through reinforcement learning and robust motor execution through rapid latent control. This dual architecture achieves significantly faster adaptation with low online computational cost compared to model-based RL baselines, while maintaining near-optimal performance. The approach combines the benefits of flexible policy learning through reinforcement learning with rapid error correction capabilities, providing a principled approach to maintaining performance in high-dimensional continuous control tasks under varying dynamics.

控制世界模型强化学习快速适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。