arXiv:2505.20962cs.ROcs.CV2025-05

用物体为中心的视觉表征提升机器人视觉运动策略学习效果

Object-Centric Action-Enhanced Representations for Robot Visuo-Motor Policy Learning

  • 将语义分割与视觉编码耦合,通过槽注意力机制实现
  • 在仿真任务中显著提升强化学习与模仿学习性能
  • 利用外部预训练模型+人类动作数据微调,减少对标注数据依赖

从观察动作中学习视觉表征以支持机器人视觉运动策略生成,是一种贴近人类认知与感知能力的有前景方向。受心理学理论启发,人类以物体为中心处理场景,我们提出一种物体为中心的编码器,将语义分割与视觉表征生成联合进行,不同于以往将二者分离的做法。为此,我们采用槽注意力机制,并使用在大规模非领域数据集上预训练的SOLV模型,对其在人类动作视频数据上进行微调。通过模拟机器人任务验证,该方法可有效提升强化学习与模仿学习的训练效果,证明了联合语义分割与编码的有效性。此外,利用在非领域数据上预训练的模型有助于该过程,且在人类动作数据上微调——尽管仍属非领域数据——能显著提升性能,因与机器人任务高度相关。这些发现表明,可降低对标注或机器人专用动作数据集的依赖,且可通过已有视觉编码器加速训练并提升泛化能力。

原文摘要 · Abstract (English)

Learning visual representations from observing actions to benefit robot visuo-motor policy generation is a promising direction that closely resembles human cognitive function and perception. Motivated by this, and further inspired by psychological theories suggesting that humans process scenes in an object-based fashion, we propose an object-centric encoder that performs semantic segmentation and visual representation generation in a coupled manner, unlike other works, which treat these as separate processes. To achieve this, we leverage the Slot Attention mechanism and use the SOLV model, pretrained in large out-of-domain datasets, to bootstrap fine-tuning on human action video data. Through simulated robotic tasks, we demonstrate that visual representations can enhance reinforcement and imitation learning training, highlighting the effectiveness of our integrated approach for semantic segmentation and encoding. Furthermore, we show that exploiting models pretrained on out-of-domain datasets can benefit this process, and that fine-tuning on datasets depicting human actions -- although still out-of-domain -- , can significantly improve performance due to close alignment with robotic tasks. These findings show the capability to reduce reliance on annotated or robot-specific action datasets and the potential to build on existing visual encoders to accelerate training and improve generalizability.

机器人学习视觉表征物体中心迁移学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。