arXiv:2411.10049cs.RO2024-11

用局部几何信息指导3D场景生成多任务操作位姿,无需针对每项任务设计算法。

SPLIT: SE(3)-diffusion via Local Geometry-based Score Prediction for 3D Scene-to-Pose-Set Matching Problems

  • 基于样本位姿的局部几何预测得分,提升生成效率
  • 单模型完成杯子翻转与悬挂任务的位姿生成
  • 适用于需灵活响应多种任务的机器人感知系统

为实现机器人通用操作能力,需从原始场景中检测出不同任务所需的操作位姿。当前许多感知算法针对特定任务设计,限制了感知模块的灵活性。本文提出一种通用问题形式——3D场景到位姿集的匹配,直接从场景中匹配对应位姿,无需依赖任务特定启发式方法。为此,我们引入SPLIT,一种用于从场景生成位姿样本的SE(3)扩散模型。该模型通过基于样本位姿的局部几何信息预测得分,实现高效生成。此外,借助扩散模型的条件生成能力,我们证明SPLIT可在单一模型中生成完成杯具翻转和悬挂操作所需的多任务位姿。

原文摘要 · Abstract (English)

To enable versatile robot manipulation, robots must detect task-relevant poses for different purposes from raw scenes. Currently, many perception algorithms are designed for specific purposes, which limits the flexibility of the perception module. We present a general problem formulation called 3D scene-to-pose-set matching, which directly matches the corresponding poses from the scene without relying on task-specific heuristics. To address this, we introduce SPLIT, an SE(3)-diffusion model for generating pose samples from a scene. The model's efficiency comes from predicting scores based on local geometry with respect to the sample pose. Moreover, leveraging the conditioned generation capability of diffusion models, we demonstrate that SPLIT can generate the multi-purpose poses, required to complete both the mug reorientation and hanging manipulation within a single model.

3D生成位姿预测扩散模型机器人操作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。