从点云直接学习复杂场景中的机械臂运动可行性,提升规划效率。
Learning Motion Feasibility from Point Clouds in Cluttered Environments

- 用点云输入训练模型预测运动是否可行,绕过传统采样规划的高成本。
- 在88个真实物体、190个杂乱桌面场景上构建270万条标签数据集,验证效果。
- 最佳模型(点云Transformer)在新物体上准确率达99.6%,速度远超传统方法。
运动可行性预测在机器人领域至关重要,尤其在任务与运动规划及操作中。在杂乱环境中,基于采样的运动规划器(SBMPs)的不可行尝试会带来巨大计算开销。现有不可行性认证方法受限于低维配置空间,且常假设由已知参数的简单几何体构成的简化环境。本文研究从原始RGB-D观测中直接学习7自由度机械臂在真实杂乱场景下的运动可行性预测。我们构建首个大规模基准数据集,包含270万条抓取可行性标签,覆盖88个扫描物体和190个杂乱桌面场景。在相同训练条件下,对比了三种代表性分类器族:基于MLP、体素卷积神经网络和点云Transformer架构。最佳模型GRASPFC-PTX(点云Transformer)在新物体上的AUROC达到0.996,预测速度显著快于SBMPs。
原文摘要 · Abstract (English)
Motion feasibility prediction plays a central role in robotics, particularly in task and motion planning and manipulation. A major bottleneck for this problem in cluttered environments is that infeasible planning attempts by Sampling-based motion planners (SBMPs) can incur substantial computational cost. Also existing approaches for infeasibility certification are limited to low-dimensional configuration spaces and often assume simplified geometric environments represented by primitive objects with known parameters. We study the complementary problem of learning motion feasibility prediction directly from raw RGB-D observations for a 7-DOF manipulator operating in realistic cluttered scenes. We introduce the first large-scale benchmark for this setting, comprising 2.7M grasp feasibility labels over 88 scanned objects and 190 cluttered tabletop scenes. We benchmark three representative classifier families spanning MLP- based, volumetric-CNN, and point-cloud-based Transformer architectures under matched training conditions. Our best model, GRASPFC-PTX (a point-cloud transformer), achieves an AUROC of 0.996 on Novel objects while providing predictions significantly faster than SBMPs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。