让手部动作成为物体位姿追踪的互补线索,提升遮挡下的追踪精度。
ComPose: When to Trust Hands for Object Pose Tracking

- 将手部运动视为互补线索,融合手与物的视觉特征进行联合追踪
- 在严重遮挡下仍保持稳定3D轨迹,旋转与平移均具时间一致性
- 适合需要从视频中重建人类操作动作的机器人应用
从视频中重建物体运动是具身智能与机器人操作的关键。现有方法依赖深度数据或3D模板等强先验,且在手部遮挡下仍易失效。本文提出ComPose,一种基于RGB视频的6自由度手感知物体位姿追踪框架。不将手仅视作遮挡物,而是将其运动作为互补线索。通过统一追踪流程,融合基础模型提取的物体与手部特征,自适应选择有效手部关节点,结合二者线索估计运动,并利用可见几何证据和学习修正优化结果。同时强化旋转与平移的时间一致性,无需外部平滑即可获得稳定3D轨迹。大量实验表明,该方法在严重遮挡与几何模糊下仍具高精度、高效性与鲁棒性。生成的轨迹可直接用于下游机器人操作,实现从在线视频中重建人类动作。
原文摘要 · Abstract (English)
Reconstructing the motion of objects from videos is a key component for embodied AI and robot manipulation. While diverse approaches to object pose tracking have been studied, they rely heavily on strong external priors, such as depth data or 3D templates, and remain highly vulnerable to severe occlusions by hand grasps despite the use of explicit masks. In this work, we present ComPose, a 6DoF object tracking framework designed for hand-aware object pose estimation from RGB video. Rather than treating the hand purely as an occluder, our method harmonizes hand motions as a \textit{complementary cue} for object tracking. In detail, we recover a variety of object motions over time by combining object and hand cues from foundation models within a unified tracking pipeline. Here, ComPose adaptively selects informative hand joints, combines object- and hand-derived cues for motion estimation, and refines the resulting object motion using visible geometric evidence and a learned correction. We further enforce the temporal consistency over both rotation and translation, yielding stable 3D object trajectories over time without any external smoothing. Extensive experiments show that our method is accurate, efficient, and robust under severe hand occlusion and geometric ambiguity. In addition, the resulting trajectories can also effectively transfer to downstream robot manipulation by enabling robots to reconstruct human actions from online videos.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。