用预训练视觉编码器缓解仿真到现实的迁移难题
Bridging the Sim2Real Gap: Vision Encoder Pre-Training for Visuomotor Policy Transfer
- 用大规模预训练视觉编码器提取机器人控制所需特征
- 23个编码器中,操作预训练模型表现最佳,CNN比ViT更稳定
- 提出两个评估指标,指导真实场景下的策略迁移
仿真为学习视觉-运动机器人策略提供了可扩展且高效的替代方案。然而,仿真到现实(Sim2Real)的分布偏移——即在真实环境中使用仿真训练的策略——常导致策略迁移失败。本文提出一种离线框架,评估大规模预训练视觉编码器缓解Sim2Real差距的能力。我们考察了多样化的编码器,评估其提取机器人控制所需特征的能力(动作得分,Action Score)以及对任务无关环境变化的不变性(领域不变性得分,Domain Invariance Score)。通过对23个编码器的评估,揭示了架构、预训练数据集和参数规模的影响规律:操作预训练编码器在动作得分上持续领先,基于CNN的编码器在领域不变性上优于ViT,而性能最佳的模型同时具备这两项特性,表明动作得分与领域不变性是互补的Sim2Real迁移能力预测因子。
原文摘要 · Abstract (English)
Simulation offers a scalable and efficient alternative to real-world data collection for learning visuomotor robotic policies. However, the simulation-to-reality, or Sim2Real distribution shift -- introduced by employing simulation-trained policies in real-world environments -- frequently prevents successful policy transfer. We present an offline framework to evaluate the performance of using large-scale pre-trained vision encoders to address the Sim2Real gap. We examine a diverse collection of encoders, assessing their ability to extract features necessary for robot control (Action Score) while remaining invariant to task-irrelevant environmental variations (Domain Invariance Score). Evaluating 23 encoders, we reveal patterns across architectures, pre-training datasets, and parameter scales. Our findings show that manipulation-pretrained encoders consistently achieve higher Action Scores, CNN-based encoders demonstrate stronger domain invariance than ViTs, and the best-performing models combine both properties, underscoring DIS and AS as complementary predictors of Sim2Real transferability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。