arXiv:2501.16389cs.ROcs.CV2025-01被引 3

用预训练视觉编码器缓解仿真到现实的迁移难题

Bridging the Sim2Real Gap: Vision Encoder Pre-Training for Visuomotor Policy Transfer

  • 用大规模预训练视觉编码器提取机器人控制所需特征
  • 23个编码器中,操作预训练模型表现最佳,CNN比ViT更稳定
  • 提出两个评估指标,指导真实场景下的策略迁移

仿真为学习视觉-运动机器人策略提供了可扩展且高效的替代方案。然而,仿真到现实(Sim2Real)的分布偏移——即在真实环境中使用仿真训练的策略——常导致策略迁移失败。本文提出一种离线框架,评估大规模预训练视觉编码器缓解Sim2Real差距的能力。我们考察了多样化的编码器,评估其提取机器人控制所需特征的能力(动作得分,Action Score)以及对任务无关环境变化的不变性(领域不变性得分,Domain Invariance Score)。通过对23个编码器的评估,揭示了架构、预训练数据集和参数规模的影响规律:操作预训练编码器在动作得分上持续领先,基于CNN的编码器在领域不变性上优于ViT,而性能最佳的模型同时具备这两项特性,表明动作得分与领域不变性是互补的Sim2Real迁移能力预测因子。

原文摘要 · Abstract (English)

Simulation offers a scalable and efficient alternative to real-world data collection for learning visuomotor robotic policies. However, the simulation-to-reality, or Sim2Real distribution shift -- introduced by employing simulation-trained policies in real-world environments -- frequently prevents successful policy transfer. We present an offline framework to evaluate the performance of using large-scale pre-trained vision encoders to address the Sim2Real gap. We examine a diverse collection of encoders, assessing their ability to extract features necessary for robot control (Action Score) while remaining invariant to task-irrelevant environmental variations (Domain Invariance Score). Evaluating 23 encoders, we reveal patterns across architectures, pre-training datasets, and parameter scales. Our findings show that manipulation-pretrained encoders consistently achieve higher Action Scores, CNN-based encoders demonstrate stronger domain invariance than ViTs, and the best-performing models combine both properties, underscoring DIS and AS as complementary predictors of Sim2Real transferability.

Sim2Real视觉编码器机器人控制预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。