统一建模手部运动生成与估计,支持遮挡和不完整输入。
UniHand: A Unified Model for Diverse Controlled 4D Hand Motion Modeling
- 用统一扩散框架将生成与估计视为条件动作合成。
- 在严重遮挡和序列缺失下仍保持高精度,优于现有方法。
- 适合需要鲁棒手部动作建模的交互系统开发人员。
手部运动在人机交互中至关重要,但建模真实的4D手部运动(即随时间变化的3D手部姿态序列)仍具挑战。现有研究通常分为两类:(1) 估计方法从视觉观测重建精确动作,但在手部遮挡或缺失时表现不佳;(2) 生成方法利用生成先验,在多模态结构化输入下合成手部姿态,并补全不完整序列。然而这种分离限制了异构条件信号的有效利用,也阻碍了任务间知识迁移。本文提出UniHand,一个基于扩散的统一框架,将估计与生成统一为条件动作合成问题。UniHand通过联合变分自编码器将异构输入(如MANO参数、2D骨骼)嵌入共享潜在空间,实现条件对齐。视觉信息由冻结的视觉主干网络编码,专用手部感知器直接从图像特征提取手部特异性线索,无需复杂的检测与裁剪流程。潜在扩散模型则基于这些多样化条件合成一致的动作序列。在多个基准上的大量实验表明,UniHand在严重遮挡和时间不完整输入下仍能实现稳健且准确的手部动作建模。
原文摘要 · Abstract (English)
Hand motion plays a central role in human interaction, yet modeling realistic 4D hand motion (i.e., 3D hand pose sequences over time) remains challenging. Research in this area is typically divided into two tasks: (1) Estimation approaches reconstruct precise motion from visual observations, but often fail under hand occlusion or absence; (2) Generation approaches focus on synthesizing hand poses by exploiting generative priors under multi-modal structured inputs and infilling motion from incomplete sequences. However, this separation not only limits the effective use of heterogeneous condition signals that frequently arise in practice, but also prevents knowledge transfer between the two tasks. We present UniHand, a unified diffusion-based framework that formulates both estimation and generation as conditional motion synthesis. UniHand integrates heterogeneous inputs by embedding structured signals into a shared latent space through a joint variational autoencoder, which aligns conditions such as MANO parameters and 2D skeletons. Visual observations are encoded with a frozen vision backbone, while a dedicated hand perceptron extracts hand-specific cues directly from image features, removing the need for complex detection and cropping pipelines. A latent diffusion model then synthesizes consistent motion sequences from these diverse conditions. Extensive experiments across multiple benchmarks demonstrate that UniHand delivers robust and accurate hand motion modeling, maintaining performance under severe occlusions and temporally incomplete inputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。