arXiv:2601.01050cs.CVcs.AI2026-01被引 8

从第一视角视频中重建手物交互的全局空间关系,支持任意物体类别。

EgoGrasp: World-Space Hand-Object Interaction Estimation from Egocentric Videos

  • 分三阶段处理:先重建场景,再用身体先验估计手姿,最后融合交互约束补全6自由度姿态。
  • 在多个开放词汇物体上实现当前最优的全局手物交互重建性能。
  • 特别适合需要理解日常动作的机器人、虚拟现实等应用。

我们提出EgoGrasp,首个从动态第一视角视频中重建世界空间手物交互(W-HOI)的方法,支持开放词汇物体。准确的W-HOI重建对具身智能至关重要,但现有方法大多局限于局部相机坐标或单帧,无法捕捉全局时序动态。尽管部分近期工作尝试世界空间手部估计,却忽略了物体位姿与手物交互约束。此外,以往方法要么因依赖物体模板而无法处理开放集类别,要么采用可微渲染需逐实例优化,计算成本过高。第一视角视频中频繁遮挡也严重降低性能。为此,我们提出多阶段框架:(i) 利用视觉基础模型进行初始3D场景、手部与物体重建的鲁棒预处理;(ii) 基于身体引导的扩散模型,引入显式第一视角身体先验以估计手部姿态;(iii) 基于交互先验的扩散模型,实现手感知的6DoF位姿补全,确保物理合理且时序一致的W-HOI估计。实验表明,EgoGrasp在多物体与开放词汇场景下均达到当前最优的重建表现。

原文摘要 · Abstract (English)

We propose EgoGrasp, the first method to reconstruct world-space hand-object interactions (W-HOI) from dynamic egoview videos, supporting open-vocabulary objects. Accurate W-HOI reconstruction is critical for embodied intelligence yet remains challenging. Existing HOI methods are largely restricted to local camera coordinates or single frames, failing to capture global temporal dynamics. While some recent approaches attempt world-space hand estimation, they overlook object poses and HOI constraints. Moreover, previous HOI estimation methods either fail to handle open-set categories due to their reliance on object templates or employ differentiable rendering that requires per-instance optimization, resulting in prohibitive computational costs. Finally, frequent occlusions in egocentric videos severely degrade performance. To overcome these challenges, we propose a multi-stage framework: (i) a robust pre-processing pipeline leveraging vision foundation models for initial 3D scene, hand and object reconstruction; (ii) a body-guided diffusion model that incorporates explicit egocentric body priors for hand pose estimation; and (iii) an HOI-prior-informed diffusion model for hand-aware 6DoF pose infilling, ensuring physically plausible and temporally consistent W-HOI estimation. We experimentally demonstrate that EgoGrasp can achieve state-of-the-art performance in W-HOI reconstruction, handling multiple and open vocabulary objects robustly.

手物交互第一视角扩散模型3D重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。