arXiv:2412.04244cs.CV2024-12CVPR被引 46

构建超大规模双手动作数据集,助力机器人与AI理解复杂手部行为

GigaHands: A Massive Annotated Dataset of Bimanual Hand Activities

  • 采用无标记捕捉技术自动获取3D手物姿态,降低标注成本
  • 覆盖56人、417种物体,含14000段动作片段和8.4万条文本注释
  • 适用于文本驱动动作生成、手部动作描述等前沿应用

理解双手人类手部活动是人工智能与机器人领域的重要挑战。现有数据集规模小、覆盖动作类型有限且标注不详尽,制约了大规模模型的发展。本文提出GigaHands,一个大规模标注数据集,包含56名受试者在417种物体上的34小时双手动作,共生成1.83亿帧图像与1.4万段动作片段,配套8.4万条文本标注。通过无标记捕捉系统与自动化数据采集流程,实现3D手部与物体姿态的全自动估计,大幅减少人工标注负担。其规模与多样性支持多种应用场景,包括文本驱动的动作合成、手部运动描述生成及动态辐射场重建。相关资源详见https://ivl.cs.brown.edu/research/gigahands.html。

原文摘要 · Abstract (English)

Understanding bimanual human hand activities is a critical problem in AI and robotics. We cannot build large models of bimanual activities because existing datasets lack the scale, coverage of diverse hand activities, and detailed annotations. We introduce GigaHands, a massive annotated dataset capturing 34 hours of bimanual hand activities from 56 subjects and 417 objects, totaling 14k motion clips derived from 183 million frames paired with 84k text annotations. Our markerless capture setup and data acquisition protocol enable fully automatic 3D hand and object estimation while minimizing the effort required for text annotation. The scale and diversity of GigaHands enable broad applications, including text-driven action synthesis, hand motion captioning, and dynamic radiance field reconstruction. Our website are avaliable at https://ivl.cs.brown.edu/research/gigahands.html .

动作识别数据集双手交互3D姿态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。