用手腕运动与相机朝向的物理耦合关系,实现无需场景几何的自指相机方向估计。
WristCompass: Kinematic Coupling as a Learnable Visual Concept for Ego-Camera Orientation

- 利用手臂链的物理结构建立手腕运动与相机朝向的动态关联
- 在桌面上训练后零样本迁移至厨房视频,中位几何误差14.3°
- 仅需200K参数的GRU即可接近10亿参数模型性能
从操作视频中恢复自指相机朝向是分离手部运动与相机运动的关键,对基于第一人称示范的模仿学习至关重要。传统依赖场景几何的方法在手部遮挡时失效:拥有10亿参数的VGGT模型在TACO基准上表现甚至不如常数预测器。我们发现一种在场景几何缺失时依然存在的替代视觉概念——运动学耦合动力学,即由臂-肩-头链结构施加的手腕运动与相机朝向之间的有序物理关系。该概念具有紧凑性(4维手腕特征优于126维全手关节点)、时序性(需短窗内GRU而非逐帧检索)和物理根基性(因源于解剖结构而可零样本跨数据集迁移)。仅在桌面操作数据上训练的WristCompass,零样本迁移至Epic Kitchens烹饪视频,达到14.3°中位几何误差,仅用200K GRU参数即逼近10亿参数场景模型性能。
原文摘要 · Abstract (English)
Recovering ego-camera orientation from manipulation video is a prerequisite for disentangling hand motion from camera motion, a key step in imitation learning from egocentric demonstrations. The obvious approach, inferring orientation from scene geometry, fails when hands occlude the frame: VGGT, a 1B-parameter scene reconstruction model, scores worse than a constant predictor on the TACO benchmark. We identify an alternative visual concept that is present precisely when scene geometry is absent: kinematic coupling dynamics, the structured physical relationship between wrist motion and camera orientation imposed by the arm-shoulder-head chain. We find that this concept is compact (4D inter-wrist features outperform 126D full hand keypoints), temporal (requiring a GRU over short windows rather than per-frame retrieval), and physically grounded (transferring zero-shot across datasets because it is rooted in anatomy rather than scene appearance). Trained only on tabletop manipulation, WristCompass transfers zero-shot to Epic Kitchens cooking video, achieving 14.3$^\circ$ median geodesic error and approaching the performance of a 1B-parameter scene model at 200K GRU parameters.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。