arXiv:2508.01600cs.RO2025-08被引 11

用动作序列对比学习提升机器人抓取的跨场景泛化能力

CLASS: Contrastive Learning via Action Sequence Supervision for Robot Manipulation

  • 通过动态时间规整识别相似动作序列,构建对比学习正样本对
  • 在5个仿真和3个真实任务中实现75%平均成功率,显著优于基线
  • 适合需要强泛化能力的机器人操控场景,尤其视觉变化大的情况

行为克隆(BC)在机器人操作中表现优异,得益于强大模型、动作序列建模和大规模示范数据。然而,在异构数据集上(如相机视角或物体外观变化),性能会显著下降,因BC容易过拟合单个示范而非捕捉共享结构,限制泛化能力。为此,我们提出基于动作序列监督的对比学习方法(CLASS),利用动态时间规整(DTW)识别相似动作序列作为弱监督信号,通过相似度加权的软InfoNCE损失优化表示。在5个仿真基准和3个真实任务上评估,仅用预训练表示即可实现竞争性检索控制结果。尤其在严重视觉偏移下,使用CLASS预训练的扩散策略平均成功率达75%,而所有基线方法均无法有效表现。

原文摘要 · Abstract (English)

Recent advances in Behavior Cloning (BC) have led to strong performance in robotic manipulation, driven by expressive models, sequence modeling of actions, and large-scale demonstration data. However, BC faces significant challenges when applied to heterogeneous datasets, such as visual shift with different camera poses or object appearances, where performance degrades despite the benefits of learning at scale. This stems from BC's tendency to overfit individual demonstrations rather than capture shared structure, limiting generalization. To address this, we introduce Contrastive Learning via Action Sequence Supervision (CLASS), a method for learning behavioral representations from demonstrations using supervised contrastive learning. CLASS leverages weak supervision from similar action sequences identified via Dynamic Time Warping (DTW) and optimizes a soft InfoNCE loss with similarity-weighted positive pairs. We evaluate CLASS on 5 simulation benchmarks and 3 real-world tasks to achieve competitive results using retrieval-based control with representations only. Most notably, for downstream policy learning under significant visual shifts, Diffusion Policy with CLASS pre-training achieves an average success rate of 75%, while all other baseline methods fail to perform competitively. Project webpage: https://class-robot.github.io.

机器人操控对比学习动作序列泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。