arXiv:2506.10359cs.ROcs.LG2025-06中稿 · Robotics: Science …被引 5

用真实数据训练多模态模型,提升机器人抓取多样性物品的准确率。

Demonstrating Multi-Suction Item Picking at Scale via Multi-Modal Learning of Pick Success

  • 融合RGB、深度和语义分割多模态信息预测抓取成功率
  • 在多个真实场景数据集上实现超过90%的抓取成功率
  • 适合工业级自动化仓库的高通量抓取系统研发

本研究展示了如何从部署于工业规模的真实世界稀疏标注数据中自主学习机器人操作特性,从而获得性能更优的解决方案。聚焦于多吸盘机器人抓取任务,全面研究了多模态视觉编码器在预测候选抓取成功概率中的应用。从杂乱堆叠中抓取多样化物品是仓储等现实场景中机器人操作的关键挑战,要求方法能适应开放集合的物品,同时满足低延迟以实现高吞吐。所提方法利用RGB、深度图和语义分割等多种输入模态,评估多吸盘抓取质量。模型基于真实抓取数据,采用多模态预训练与微调相结合的方式进行训练。论文在大规模物品抓取数据集、含部分遮挡的物品抓取数据集及面向容器(如箱子、信封)的包装件抓取数据集上进行了综合实验,评估了不同物品配置、抓取场景和物体类型下的表现。消融实验揭示了域内预训练的重要性,不同模态的影响以及微调的必要性;结果表明,模型能在预训练阶段学习模态间的关联关系,使得在微调和推理时仅需使用部分模态即可保持高性能。

原文摘要 · Abstract (English)

This work demonstrates how autonomously learning aspects of robotic operation from sparsely-labeled, real-world data of deployed, engineered solutions at industrial scale can provide with solutions that achieve improved performance. Specifically, it focuses on multi-suction robot picking and performs a comprehensive study on the application of multi-modal visual encoders for predicting the success of candidate robotic picks. Picking diverse items from unstructured piles is an important and challenging task for robot manipulation in real-world settings, such as warehouses. Methods for picking from clutter must work for an open set of items while simultaneously meeting latency constraints to achieve high throughput. The demonstrated approach utilizes multiple input modalities, such as RGB, depth and semantic segmentation, to estimate the quality of candidate multi-suction picks. The strategy is trained from real-world item picking data, with a combination of multimodal pretrain and finetune. The manuscript provides comprehensive experimental evaluation performed over a large item-picking dataset, an item-picking dataset targeted to include partial occlusions, and a package-picking dataset, which focuses on containers, such as boxes and envelopes, instead of unpackaged items. The evaluation measures performance for different item configurations, pick scenes, and object types. Ablations help to understand the effects of in-domain pretraining, the impact of different modalities and the importance of finetuning. These ablations reveal both the importance of training over multiple modalities but also the ability of models to learn during pretraining the relationship between modalities so that during finetuning and inference, only a subset of them can be used as input.

机器人抓取多模态学习工业自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。