预训练视觉模型让机器人在视觉变化时更稳定,尤其适合复杂环境中的模型预测控制。
Pre-trained Visual Representations Generalize Where it Matters in Model-Based Reinforcement Learning
- 用预训练视觉模型提升策略网络,减少对新场景的敏感度
- 极端视觉变化下,部分微调的模型性能比从零训练高30%以上
- 适合做视觉感知鲁棒性要求高的机器人控制任务
在视觉运动策略学习中,机器人直接从视觉输入生成控制指令。传统方法联合训练策略与视觉编码器,对新视觉场景泛化能力差。使用预训练视觉模型(PVMs)可提升无模型强化学习的鲁棒性。尽管模型基础强化学习(MBRL)更高效,但现有研究发现PVMs在MBRL中效果不佳。本文研究了PVM在MBRL中对视觉域偏移下的泛化能力。结果表明,在严重视觉变化场景下,使用PVM的模型显著优于从零训练的基线。进一步实验显示,适度微调可在最极端分布偏移下保持最高平均任务性能。这证明预训练视觉模型能有效增强视觉策略学习的鲁棒性,为模型基机器人学习广泛应用提供了有力证据。
原文摘要 · Abstract (English)
In visuomotor policy learning, the control policy for the robotic agent is derived directly from visual inputs. The typical approach, where a policy and vision encoder are trained jointly from scratch, generalizes poorly to novel visual scene changes. Using pre-trained vision models (PVMs) to inform a policy network improves robustness in model-free reinforcement learning (MFRL). Recent developments in Model-based reinforcement learning (MBRL) suggest that MBRL is more sample-efficient than MFRL. However, counterintuitively, existing work has found PVMs to be ineffective in MBRL. Here, we investigate PVM's effectiveness in MBRL, specifically on generalization under visual domain shifts. We show that, in scenarios with severe shifts, PVMs perform much better than a baseline model trained from scratch. We further investigate the effects of varying levels of fine-tuning of PVMs. Our results show that partial fine-tuning can maintain the highest average task performance under the most extreme distribution shifts. Our results demonstrate that PVMs are highly successful in promoting robustness in visual policy learning, providing compelling evidence for their wider adoption in model-based robotic learning applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。