arXiv:2409.20248cs.RO2024-09ICRA被引 1

视觉编码器不只是特征提取,还能参与决策,影响机器人控制效果。

Feature Extractor or Decision Maker: Rethinking the Role of Visual Encoders in Visuomotor Policies

  • 通过视觉对齐测试,发现端到端模型中编码器主动参与决策。
  • 使用域外预训练的编码器性能平均下降42%,远低于端到端模型。
  • 适合关注视觉编码器作用、改进预训练方法的研究者阅读。

端到端视觉-运动策略通常被视为一个整体,但近期采用域外(OOD)数据预训练视觉编码器的方法,将编码器与后续网络分离,后者被称为策略。我们提出视觉对齐测试(Visual Alignment Testing),用于评估这种功能分离的有效性。结果表明,在端到端训练的模型中,视觉编码器因受到运动数据监督而积极参与决策,与假设的功能分离相矛盾。相比之下,未经过运动监督的域外预训练编码器在基准测试中平均性能下降42%,显著低于当前最优的端到端策略表现。我们认为,这一初步探索有助于未来预训练方法的发展,例如设计任务感知或上下文感知的编码器。

原文摘要 · Abstract (English)

An end-to-end (E2E) visuomotor policy is typically treated as a unified whole, but recent approaches using out-of-domain (OOD) data to pretrain the visual encoder have cleanly separated the visual encoder from the network, with the remainder referred to as the policy. We propose Visual Alignment Testing, an experimental framework designed to evaluate the validity of this functional separation. Our results indicate that in E2E-trained models, visual encoders actively contribute to decision-making resulting from motor data supervision, contradicting the assumed functional separation. In contrast, OOD-pretrained models, where encoders lack this capability, experience an average performance drop of 42\% in our benchmark results, compared to the state-of-the-art performance achieved by E2E policies. We believe this initial exploration of visual encoders' role can provide a first step towards guiding future pretraining methods to address their decision-making ability, such as developing task-conditioned or context-aware encoders.

视觉编码器机器人控制端到端学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。