构建首个完整抓握分类的真实3D手物交互数据集
Dense Hand-Object(HO) GraspNet with Full Grasping Taxonomy and Dynamics
- 基于抓握分类定义原子动作,组合生成复杂交互行为
- 涵盖22种刚性物+8种复合物,99人不同年龄,150万帧带标注
- 支持抓握分类与3D手姿估计,适合构建通用手物交互模型
现有3D手物交互数据集在数据量、交互场景多样性或标注质量上均受限。本文提出全新的训练数据集HOGraspNet,是唯一真实采集并覆盖完整抓握分类的数据集,包含丰富的类内差异。以抓握分类作为基本动作单元,其时空组合可表征围绕物体的复杂手部活动。从YCB数据集中选取22个刚性物体及8个其他复合物体,依据形状与尺寸分类,确保覆盖所有手部抓握形态。数据集包含99名年龄10至74岁参与者的多样化手形,连续视频帧,以及150万帧稀疏的RGB-D数据和标注。提供3D手与物体网格、3D关键点、接触图及抓握标签。通过多视角RGB-D帧拟合手部参数化模型(MANO)和隐式函数(HALO),获得精确3D网格;仅对物体使用动作捕捉系统。其中HALO拟合无需参数调优,扩展性好,精度与MANO相当。在抓握分类与3D手姿估计任务上评估,结果显示性能随抓握类型与物体类别变化,表明数据集所捕获的交互空间具有重要价值。该数据旨在支持学习通用形状先验或3D手物交互基础模型。数据与代码已公开于https://hograspnet2024.github.io/。
原文摘要 · Abstract (English)
Existing datasets for 3D hand-object interaction are limited either in the data cardinality, data variations in interaction scenarios, or the quality of annotations. In this work, we present a comprehensive new training dataset for hand-object interaction called HOGraspNet. It is the only real dataset that captures full grasp taxonomies, providing grasp annotation and wide intraclass variations. Using grasp taxonomies as atomic actions, their space and time combinatorial can represent complex hand activities around objects. We select 22 rigid objects from the YCB dataset and 8 other compound objects using shape and size taxonomies, ensuring coverage of all hand grasp configurations. The dataset includes diverse hand shapes from 99 participants aged 10 to 74, continuous video frames, and a 1.5M RGB-Depth of sparse frames with annotations. It offers labels for 3D hand and object meshes, 3D keypoints, contact maps, and \emph{grasp labels}. Accurate hand and object 3D meshes are obtained by fitting the hand parametric model (MANO) and the hand implicit function (HALO) to multi-view RGBD frames, with the MoCap system only for objects. Note that HALO fitting does not require any parameter tuning, enabling scalability to the dataset's size with comparable accuracy to MANO. We evaluate HOGraspNet on relevant tasks: grasp classification and 3D hand pose estimation. The result shows performance variations based on grasp type and object class, indicating the potential importance of the interaction space captured by our dataset. The provided data aims at learning universal shape priors or foundation models for 3D hand-object interaction. Our dataset and code are available at https://hograspnet2024.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。