从单目视频中预测连贯的4D手部运动轨迹
Predicting 4D Hand Trajectory from Monocular Videos
- 用多帧输入的Transformer直接预测连贯4D轨迹
- 在全局轨迹精度上显著优于现有方法
- 适合需要精准手部运动建模的应用场景
我们提出HaPTIC,一种从单目视频中推断连贯4D手部轨迹的方法。当前基于视频的手部姿态重建方法主要关注相邻帧间的帧内3D姿态优化,而非空间中的连续4D轨迹。尽管具有时间线索,但因标注视频数据稀缺,性能通常低于基于图像的方法。为此,我们复用先进的图像级Transformer,输入多帧并直接预测连贯轨迹。引入两种轻量级注意力层:跨视图自注意力用于融合时序信息,全局交叉注意力用于引入更大空间上下文。该方法生成的4D轨迹与真实轨迹高度一致,同时保持强2D重投影对齐。我们在第一人称和第三人称视频上均验证了有效性,在全局轨迹精度上显著超越现有方法,且在单图姿态估计上达到最先进的水平。
原文摘要 · Abstract (English)
We present HaPTIC, an approach that infers coherent 4D hand trajectories from monocular videos. Current video-based hand pose reconstruction methods primarily focus on improving frame-wise 3D pose using adjacent frames rather than studying consistent 4D hand trajectories in space. Despite the additional temporal cues, they generally underperform compared to image-based methods due to the scarcity of annotated video data. To address these issues, we repurpose a state-of-the-art image-based transformer to take in multiple frames and directly predict a coherent trajectory. We introduce two types of lightweight attention layers: cross-view self-attention to fuse temporal information, and global cross-attention to bring in larger spatial context. Our method infers 4D hand trajectories similar to the ground truth while maintaining strong 2D reprojection alignment. We apply the method to both egocentric and allocentric videos. It significantly outperforms existing methods in global trajectory accuracy while being comparable to the state-of-the-art in single-image pose estimation. Project website: https://judyye.github.io/haptic-www
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。