从局部视图重建物体三维结构并生成任务导向抓取策略
Volumetric Reconstruction From Partial Views for Task-Oriented Grasping
- 用带LSTM的循环生成对抗网络处理不定数量深度扫描
- 在四类任务中实现89%的抓取准确率
- 结合先验知识与强化学习优化抓取动作,适合机器人操作场景
物体可操作性与体积信息对制定任务约束下的有效抓取策略至关重要。本文提出一种从有限局部视图推断合适抓取策略的方法。通过引入带有长短期记忆(LSTM)单元的循环生成对抗网络(R-GAN),使模型能够处理任意数量的深度扫描。为确定物体可操作性,利用AffordPose知识数据集作为先验知识,通过切比雪夫距离度量体积相似性及动作相似性来检索可操作性。进一步采用近端策略优化(PPO)强化学习模型对检索到的抓取策略进行优化,以实现任务导向抓取。在双臂移动操作机器人上对四种任务(提升、手柄抓取、包裹抓取、按压)进行了评估,整体抓取准确率达到89%。
原文摘要 · Abstract (English)
Object affordance and volumetric information are essential in devising effective grasping strategies under task-specific constraints. This paper presents an approach for inferring suitable grasping strategies from limited partial views of an object. To achieve this, a recurrent generative adversarial network (R-GAN) was proposed by incorporating a recurrent generator with long short-term memory (LSTM) units for it to process a variable number of depth scans. To determine object affordances, the AffordPose knowledge dataset is utilized as prior knowledge. Affordance retrieving is defined by the volume similarity measured via Chamfer Distance and action similarities. A Proximal Policy Optimization (PPO) reinforcement learning model is further implemented to refine the retrieved grasp strategies for task-oriented grasping. The retrieved grasp strategies were evaluated on a dual-arm mobile manipulation robot with an overall grasping accuracy of 89% for four tasks: lift, handle grasp, wrap grasp, and press.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。