arXiv:2602.22209cs.CV2026-02被引 7

从第一视角视频中联合重建手与物体的运动轨迹

WHOLE: World-Grounded Hand-Object Lifted from Egocentric Videos

  • 基于生成先验联合建模手物交互关系
  • 在手姿、物体位姿和相对关系上均达新最优
  • 适合需要精准动作理解的智能机器人场景

第一视角操作视频因交互过程中的严重遮挡以及物体频繁进出视野而极具挑战。现有方法通常孤立恢复手或物体姿态,二者在交互时均表现不佳,且难以处理视线外情况,独立预测常导致手物关系不一致。本文提出WHOLE,一种从第一视角视频中结合物体模板重建手与物体世界空间运动的方法。核心思想是学习手物运动的生成先验,以联合推理其交互关系。测试时,预训练先验被引导生成符合视频观测的轨迹。该联合生成重建显著优于先分别处理再后处理的方法。WHOLE在手部运动估计、6D物体位姿估计及相对交互重建任务上均达到当前最优性能。

原文摘要 · Abstract (English)

Egocentric manipulation videos are highly challenging due to severe occlusions during interactions and frequent object entries and exits from the camera view as the person moves. Current methods typically focus on recovering either hand or object pose in isolation, but both struggle during interactions and fail to handle out-of-sight cases. Moreover, their independent predictions often lead to inconsistent hand-object relations. We introduce WHOLE, a method that holistically reconstructs hand and object motion in world space from egocentric videos given object templates. Our key insight is to learn a generative prior over hand-object motion to jointly reason about their interactions. At test time, the pretrained prior is guided to generate trajectories that conform to the video observations. This joint generative reconstruction substantially outperforms approaches that process hands and objects separately followed by post-processing. WHOLE achieves state-of-the-art performance on hand motion estimation, 6D object pose estimation, and their relative interaction reconstruction. Project website: https://judyye.github.io/whole-www

动作重建第一视角手物交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。