用隐状态残差修正仿真模型,提升视觉强化学习的现实部署效果
Adapting World Models with Latent-State Dynamics Residuals
- 在隐空间中对动态模型做残差修正,避免高维图像直接校准
- 在少数据场景下仍能有效建模真实动态,性能优于传统迁移方法
- 适合视觉输入的机器人任务,尤其适用于仿真到现实的迁移
仿真到现实的强化学习面临仿真与真实动态差异的挑战,严重降低智能体性能。现有方法通过学习模拟器前向动态的残差误差函数进行修正,但在高维状态(如图像)下不适用。为此,我们提出ReDRAW:一种在仿真中预训练的隐状态自回归世界模型,通过修正隐状态动态而非显式观测状态来适配目标环境。利用该适配后的世界模型,ReDRAW使强化学习智能体能在修正后的动态下进行想象式滚动优化,并部署于真实世界。在多个基于视觉的MuJoCo任务及一个物理机器人视觉车道跟随任务中,ReDRAW能有效建模动态变化,在数据稀缺时避免过拟合,而传统迁移方法在此类条件下失效。
原文摘要 · Abstract (English)
Simulation-to-reality reinforcement learning (RL) faces the critical challenge of reconciling discrepancies between simulated and real-world dynamics, which can severely degrade agent performance. A promising approach involves learning corrections to simulator forward dynamics represented as a residual error function, however this operation is impractical with high-dimensional states such as images. To overcome this, we propose ReDRAW, a latent-state autoregressive world model pretrained in simulation and calibrated to target environments through residual corrections of latent-state dynamics rather than of explicit observed states. Using this adapted world model, ReDRAW enables RL agents to be optimized with imagined rollouts under corrected dynamics and then deployed in the real world. In multiple vision-based MuJoCo domains and a physical robot visual lane-following task, ReDRAW effectively models changes to dynamics and avoids overfitting in low data regimes where traditional transfer methods fail.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。