通过多层级对齐表征,提升机器人操作任务成功率预测精度
Task Success Prediction for Open-Vocabulary Manipulation Based on Multi-Level Aligned Representations

- 融合图像局部特征、语言对齐特征与语言结构特征构建多级表征
- 在真实机器人平台测试中,准确率比主流多模态大模型高8.66点
- 适合需要精准判断操作效果的具身智能与人机协作场景
本研究针对基于指令文本和第一人称视角图像的开放词汇操作任务成功预测问题,提出对比性λ-Repformer模型。传统方法如多模态大语言模型常难以理解物体细节或位置细微变化。该方法通过整合三类特征构建多层级对齐表征:保留图像局部信息的特征、与自然语言对齐的特征、以及由自然语言结构化的特征,使模型能聚焦于两帧图像间表征差异以判断关键变化。我们在基于RT-1大规模标准数据集的自建数据集及物理机器人平台上进行评估,结果表明该方法优于现有模型,最佳模型相比代表性多模态大模型准确率提升8.66个百分点。
原文摘要 · Abstract (English)
In this study, we consider the problem of predicting task success for open-vocabulary manipulation by a manipulator, based on instruction sentences and egocentric images before and after manipulation. Conventional approaches, including multimodal large language models (MLLMs), often fail to appropriately understand detailed characteristics of objects and/or subtle changes in the position of objects. We propose Contrastive $λ$-Repformer, which predicts task success for table-top manipulation tasks by aligning images with instruction sentences. Our method integrates the following three key types of features into a multi-level aligned representation: features that preserve local image information; features aligned with natural language; and features structured through natural language. This allows the model to focus on important changes by looking at the differences in the representation between two images. We evaluate Contrastive $λ$-Repformer on a dataset based on a large-scale standard dataset, the RT-1 dataset, and on a physical robot platform. The results show that our approach outperformed existing approaches including MLLMs. Our best model achieved an improvement of 8.66 points in accuracy compared to the representative MLLM-based model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。