arXiv:2409.18297cs.ROcs.AI2024-09ICRA被引 3

构建大规模衣物操作数据集,助力机器人更好抓取折叠衣物

Flat'n'Fold: A Diverse Multi-Modal Dataset for Garment Perception and Manipulation

  • 收集1212次真人与887次机器人操作,覆盖44种衣物的展开到折叠全过程
  • 同步记录多视角RGB-D图像、点云及手/夹爪位置姿态数据,信息全面
  • 适用于机器人感知与可变形物体操作研究,推动该领域技术进步

我们提出Flat'n'Fold,一个针对衣物操作的新颖大规模数据集,填补了现有数据集的关键空白。该数据集包含1,212次人类和887次机器人对44种独特衣物在8个类别中的展开与折叠操作示范,规模、范围与多样性均超越先前基准。其独特之处在于完整捕捉从褶皱到折叠状态的整个操作过程,提供同步的多视角RGB-D图像、点云以及动作数据(包括手部或夹爪的位置与旋转)。我们量化了该数据集相较于现有基准的多样性和复杂性,表明其在视觉和动作信息上均具有真实且多样的实际操作示范。为展示数据集的实用性,我们建立了抓取点预测与子任务分解的新基准。对前沿模型的评估显示仍有显著提升空间,凸显Flat'n'Fold在推动机器人对可变形物体的感知与操作方面的潜力。数据集可通过https://cvas-ug.github.io/flat-n-fold下载。

原文摘要 · Abstract (English)

We present Flat'n'Fold, a novel large-scale dataset for garment manipulation that addresses critical gaps in existing datasets. Comprising 1,212 human and 887 robot demonstrations of flattening and folding 44 unique garments across 8 categories, Flat'n'Fold surpasses prior datasets in size, scope, and diversity. Our dataset uniquely captures the entire manipulation process from crumpled to folded states, providing synchronized multi-view RGB-D images, point clouds, and action data, including hand or gripper positions and rotations. We quantify the dataset's diversity and complexity compared to existing benchmarks and show that our dataset features natural and diverse manipulations of real-world demonstrations of human and robot demonstrations in terms of visual and action information. To showcase Flat'n'Fold's utility, we establish new benchmarks for grasping point prediction and subtask decomposition. Our evaluation of state-of-the-art models on these tasks reveals significant room for improvement. This underscores Flat'n'Fold's potential to drive advances in robotic perception and manipulation of deformable objects. Our dataset can be downloaded at https://cvas-ug.github.io/flat-n-fold

衣物操作机器人感知多模态数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。