用少量标注数据+大量视觉数据,让机器人更高效学会开门动作。
Semi-Supervised Neural Processes for Articulated Object Interactions
- 结合少量交互数据与大量无交互图像数据联合训练
- 在门开任务中性能优于其他半监督方法,仅需极少数据
- 适合标注成本高的机器人操作场景
标注动作数据的稀缺性给机器人物体操作的机器学习算法发展带来重大挑战。机器人与众多物体交互成本高昂且常不可行。相反,无需交互的物体视觉数据大量存在,可用于预训练和特征提取。然而,当前依赖图像数据预训练的方法难以适应特定任务预测,因学习到的特征未必相关。本文提出半监督神经过程(SSNP):一种适用于仅有少量物体具有标注交互数据的场景的自适应奖励预测模型。除了预测奖励标签外,SSNP的隐空间还通过更大数量物体的被动数据,以自编码目标进行联合训练。同时利用两类数据的联合训练使模型更聚焦于可泛化的特征,减少对大规模重训练的需求,从而降低计算开销。在门开任务中验证了SSNP的有效性,其表现优于其他半监督方法,且所需数据量远低于其他自适应模型。
原文摘要 · Abstract (English)
The scarcity of labeled action data poses a considerable challenge for developing machine learning algorithms for robotic object manipulation. It is expensive and often infeasible for a robot to interact with many objects. Conversely, visual data of objects, without interaction, is abundantly available and can be leveraged for pretraining and feature extraction. However, current methods that rely on image data for pretraining do not easily adapt to task-specific predictions, since the learned features are not guaranteed to be relevant. This paper introduces the Semi-Supervised Neural Process (SSNP): an adaptive reward-prediction model designed for scenarios in which only a small subset of objects have labeled interaction data. In addition to predicting reward labels, the latent-space of the SSNP is jointly trained with an autoencoding objective using passive data from a much larger set of objects. Jointly training with both types of data allows the model to focus more effectively on generalizable features and minimizes the need for extensive retraining, thereby reducing computational demands. The efficacy of SSNP is demonstrated through a door-opening task, leading to better performance than other semi-supervised methods, and only using a fraction of the data compared to other adaptive models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。