让视觉模型学会关注身体动作,提升机器人抓取效率
Grounding Bodily Awareness in Visual Representations for Efficient Policy Learning
- 用对比学习分离机器人与环境的视觉特征
- 在多种任务上提升策略性能,支持跨机器人迁移
- 适合做机器人视觉感知与策略学习的研究者
由于动作执行中复杂的机体动力学,学习有效的视觉表征仍是机器人操作中的核心挑战。本文研究具有身体相关线索的视觉表征如何促进下游机器人操作任务的高效策略学习。提出ICon(Inter-token Contrast),一种应用于视觉变换器(ViT)token级表征的对比学习方法,强制将代理特异性和环境特异性token在特征空间中分离,生成嵌入身体特定归纳偏置的代理中心视觉表征。该框架可通过引入对比损失作为辅助目标,无缝集成至端到端策略学习中。实验表明,ICon不仅在多种操作任务中提升策略表现,还促进了不同机器人间的策略迁移。
原文摘要 · Abstract (English)
Learning effective visual representations for robotic manipulation remains a fundamental challenge due to the complex body dynamics involved in action execution. In this paper, we study how visual representations that carry body-relevant cues can enable efficient policy learning for downstream robotic manipulation tasks. We present $\textbf{I}$nter-token $\textbf{Con}$trast ($\textbf{ICon}$), a contrastive learning method applied to the token-level representations of Vision Transformers (ViTs). ICon enforces a separation in the feature space between agent-specific and environment-specific tokens, resulting in agent-centric visual representations that embed body-specific inductive biases. This framework can be seamlessly integrated into end-to-end policy learning by incorporating the contrastive loss as an auxiliary objective. Our experiments show that ICon not only improves policy performance across various manipulation tasks but also facilitates policy transfer across different robots. The project website: https://inter-token-contrast.github.io/icon/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。