arXiv:2511.12030cs.CV2025-11AAAI

让手物姿态估计同时满足视觉真实与物理合理。

VPHO: Joint Visual-Physical Cue Learning and Aggregation for Hand-Object Pose Estimation

  • 联合学习视觉与物理线索,提升交互表征能力。
  • 通过扩散生成候选姿态并融合优化,精度显著提升。
  • 适合需要高真实感的虚拟现实与人机交互场景。

从单张RGB图像估计双手与物体的3D姿态是增强现实与人机交互中的基础但极具挑战性的问题。现有方法主要依赖视觉线索,常导致手物穿插或非接触等物理不成立结果。近期引入物理推理的方法多依赖后处理或不可微物理引擎,影响视觉一致性与端到端训练。为此,本文提出新框架,联合视觉与物理线索进行手物姿态估计。核心思路包括:1)联合视觉-物理线索学习:模型同步提取2D视觉线索与3D物理线索,实现对交互关系更全面的表征;2)候选姿态聚合:通过扩散模型生成多个候选姿态,结合视觉与物理预测进行融合,得到既视觉一致又物理合理的最终估计。大量实验表明,本方法在姿态精度与物理合理性上均显著优于现有最先进方法。

原文摘要 · Abstract (English)

Estimating the 3D poses of hands and objects from a single RGB image is a fundamental yet challenging problem, with broad applications in augmented reality and human-computer interaction. Existing methods largely rely on visual cues alone, often producing results that violate physical constraints such as interpenetration or non-contact. Recent efforts to incorporate physics reasoning typically depend on post-optimization or non-differentiable physics engines, which compromise visual consistency and end-to-end trainability. To overcome these limitations, we propose a novel framework that jointly integrates visual and physical cues for hand-object pose estimation. This integration is achieved through two key ideas: 1) joint visual-physical cue learning: The model is trained to extract 2D visual cues and 3D physical cues, thereby enabling more comprehensive representation learning for hand-object interactions; 2) candidate pose aggregation: A novel refinement process that aggregates multiple diffusion-generated candidate poses by leveraging both visual and physical predictions, yielding a final estimate that is visually consistent and physically plausible. Extensive experiments demonstrate that our method significantly outperforms existing state-of-the-art approaches in both pose accuracy and physical plausibility.

姿态估计物理约束扩散模型手物交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。