用视觉提示指导机器人在便利店抓取放置,提升复杂环境下的操作精度。
Visual Prompting for Robotic Manipulation with Annotation-Guided Pick-and-Place Using ACT
- 通过标注框提供抓取与放置位置的结构化空间引导
- 采用ACT模型从人类示范中学习分块动作序列,成功率显著提升
- 适合零售场景中密集、遮挡复杂的物体操作任务
便利店中的抓取放置任务因物体密集排列、遮挡以及颜色、形状、大小、纹理等属性差异而面临挑战,导致轨迹规划与抓取困难。本文提出一种感知-行动流水线,利用标注引导的视觉提示:边界框标注同时识别可抓取物体和放置位置,提供结构化空间指引。不同于传统分步规划,采用基于Transformer的动作分块(ACT)作为模仿学习算法,使机械臂能从人类示范中预测分块动作序列,实现平滑、自适应且数据驱动的抓取放置操作。我们在成功率达和抓取行为视觉分析的基础上评估系统,验证了在零售环境中抓取准确率与适应性的提升。
原文摘要 · Abstract (English)
Robotic pick-and-place tasks in convenience stores pose challenges due to dense object arrangements, occlusions, and variations in object properties such as color, shape, size, and texture. These factors complicate trajectory planning and grasping. This paper introduces a perception-action pipeline leveraging annotation-guided visual prompting, where bounding box annotations identify both pickable objects and placement locations, providing structured spatial guidance. Instead of traditional step-by-step planning, we employ Action Chunking with Transformers (ACT) as an imitation learning algorithm, enabling the robotic arm to predict chunked action sequences from human demonstrations. This facilitates smooth, adaptive, and data-driven pick-and-place operations. We evaluate our system based on success rate and visual analysis of grasping behavior, demonstrating improved grasp accuracy and adaptability in retail environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。