arXiv:2509.23888cs.CV2025-09被引 1

首个无标记3D双手与身体协同数据集,提升双臂动作识别准确率

AssemblyHands-X: Modeling 3D Hand-Body Coordination for Understanding Bimanual Human Activities

  • 通过多视角三角化+SMPL-X拟合实现无标记3D姿态标注
  • 联合建模手与身体动作使识别准确率显著优于单一模态
  • 适合研究双臂协作、人机交互等场景的动作理解任务

双臂人类活动天然涉及双手与身体的协同运动,但因缺乏合适的基准数据集,其协同作用对动作理解的影响尚未系统评估。现有3D活动数据集通常仅标注手或身体姿态,而基于标记的动捕虽可提供全身姿态,却引入视觉伪影,限制模型在自然无标记视频中的泛化能力。为此,我们提出AssemblyHands-X,首个用于双臂活动的无标记3D手-体协同基准数据集。我们构建了从同步多视角视频中进行3D姿态标注的流水线,结合多视角三角化与SMPL-X网格拟合,实现手部和上半身的可靠3D定位。我们在基于图卷积或时空注意力的近期动作识别模型上,验证了视频、手部姿态、身体姿态及手-体联合姿态等多种输入表示。实验表明,基于姿态的动作推断比视频基线更高效且更准确;联合建模手与身体线索显著提升识别性能,优于单独使用手或上半身信息,凸显建模手-体动态耦合关系对全面理解双臂活动的重要性。

原文摘要 · Abstract (English)

Bimanual human activities inherently involve coordinated movements of both hands and body. However, the impact of this coordination in activity understanding has not been systematically evaluated due to the lack of suitable datasets. Such evaluation demands kinematic-level annotations (e.g., 3D pose) for the hands and body, yet existing 3D activity datasets typically annotate either hand or body pose. Another line of work employs marker-based motion capture to provide full-body pose, but the physical markers introduce visual artifacts, thereby limiting models' generalization to natural, markerless videos. To address these limitations, we present AssemblyHands-X, the first markerless 3D hand-body benchmark for bimanual activities, designed to study the effect of hand-body coordination for action recognition. We begin by constructing a pipeline for 3D pose annotation from synchronized multi-view videos. Our approach combines multi-view triangulation with SMPL-X mesh fitting, yielding reliable 3D registration of hands and upper body. We then validate different input representations (e.g., video, hand pose, body pose, or hand-body pose) across recent action recognition models based on graph convolution or spatio-temporal attention. Our extensive experiments show that pose-based action inference is more efficient and accurate than video baselines. Moreover, joint modeling of hand and body cues improves action recognition over using hands or upper body alone, highlighting the importance of modeling interdependent hand-body dynamics for a holistic understanding of bimanual activities.

动作识别3D姿态手体协同无标记

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。