arXiv:2603.26068cs.CV2026-03中稿 · CVPR被引 2

让手部动作更符合物理规律,还能判断估计结果的可信度。

PAD-Hand: Physics-Aware Diffusion for Hand Motion Recovery

  • 用扩散模型结合物理动力学,优化手部动作序列
  • 可输出每帧每关节的物理一致性方差,量化可信度
  • 适合需要高可信手部动作估计的场景,如虚拟交互

从图像重建手部动作虽已取得显著进展,但常缺乏物理一致性,且无法衡量估计结果的可信度。本文提出一种新型物理感知的条件扩散框架,通过修正噪声姿态序列生成符合物理规律的手部运动,并估计运动估计中的物理方差。基于MeshCNN-Transformer主干网络,我们构建了刚性手部的欧拉-拉格朗日动力学模型。不同于以往强制残差为零的方法,我们将动态残差视为虚拟观测值,以更有效地融合物理约束。通过最后一层拉普拉斯近似,方法可输出每关节、每时间点的方差,用于衡量物理一致性,并提供可解释的方差图,揭示物理不一致的区域。在两个知名手部数据集上的实验表明,该方法在基于图像的初始估计上持续取得提升,且与基于视频的方法表现相当。定性结果证实,方差估计与图像基估计中动作的物理合理性高度一致。

原文摘要 · Abstract (English)

Significant advancements made in reconstructing hands from images have delivered accurate single-frame estimates, yet they often lack physics consistency and provide no notion of how confidently the motion satisfies physics. In this paper, we propose a novel physics-aware conditional diffusion framework that refines noisy pose sequences into physically plausible hand motion while estimating the physics variance in motion estimates. Building on a MeshCNN-Transformer backbone, we formulate Euler-Lagrange dynamics for articulated hands. Unlike prior works that enforce zero residuals, we treat the resulting dynamic residuals as virtual observables to more effectively integrate physics. Through a last-layer Laplace approximation, our method produces per-joint, per-time variances that measure physics consistency and offers interpretable variance maps indicating where physical consistency weakens. Experiments on two well-known hand datasets show consistent gains over strong image-based initializations and competitive video-based methods. Qualitative results confirm that our variance estimations are aligned with the physical plausibility of the motion in image-based estimates.

手部动作扩散模型物理一致性方差估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。