arXiv:2509.13349cs.ROcs.AI2025-09

用自监督预训练提升少样本抓取角度预测精度

Label-Efficient Grasp Joint Prediction with Point-JEPA

  • 基于点云的JEPA自监督预训练,支持少标注数据学习
  • 在25%数据下误差降低26%,全量数据时性能持平
  • 适合数据稀缺场景下的机器人抓取系统开发

我们研究了基于点JEPA的3D自监督预训练是否能实现标签高效的抓取关节角预测。将网格采样为点云并进行分块编码;使用ShapeNet预训练的Point-JEPA编码器,搭配K=5的多假设输出头,采用赢家通吃策略训练,并通过最高对数概率选择评估。在一个具有严格物体级划分的多指手部数据集上,点JEPA在低标签场景下显著提升了顶对数概率均方根误差(RMSE)和15°覆盖率(Coverage@15°),例如在25%数据时RMSE降低26%,在全监督条件下达到性能持平,表明JEPA式预训练是实现数据高效抓取学习的实用方法。

原文摘要 · Abstract (English)

We study whether 3D self-supervised pretraining with Point--JEPA enables label-efficient grasp joint-angle prediction. Meshes are sampled to point clouds and tokenized; a ShapeNet-pretrained Point--JEPA encoder feeds a $K{=}5$ multi-hypothesis head trained with winner-takes-all and evaluated by top--logit selection. On a multi-finger hand dataset with strict object-level splits, Point--JEPA improves top--logit RMSE and Coverage@15$^{\circ}$ in low-label regimes (e.g., 26% lower RMSE at 25% data) and reaches parity at full supervision, suggesting JEPA-style pretraining is a practical lever for data-efficient grasp learning.

自监督抓取预测少样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。