用多视角3D重建+时空变换器,解决头戴式摄像头下人体追踪难题
LAMP: Localization Aware Multi-camera People Tracking in Metric 3D World

- 先将多视角2D关键点转为统一3D世界坐标,再用变换器拟合3D运动轨迹
- 在动态头戴场景中比现有方法提升显著,尤其在遮挡和剧烈运动时更鲁棒
- 适合需要高精度3D人体追踪的智能穿戴、虚拟现实等应用
从第一人称多摄像头头戴设备中追踪3D人体运动面临严重自我运动、部分可见性或遮挡以及训练数据匮乏的挑战。现有单目视频方法通常依赖静态或缓慢移动的摄像头,无法高效利用多视角、校准且定位准确的输入,导致在动态第一人称捕捉中表现脆弱。本文提出LAMP(Localization Aware Multi-camera People Tracking):一种新颖而简单的框架,通过早期解耦观察者与目标运动来解决该问题。LAMP采用两步流程:首先,利用已知设备6自由度运动和校准信息,将所有相机在时间窗口内的检测到的2D人体关键点转换到统一的3D世界参考帧;其次,使用端到端训练的时空变换器直接拟合该3D射线云中的3D人体运动。这种“先升维再拟合”的方法使LAMP能够学习并利用世界空间中的自然人体运动先验,同时提供一个灵活框架,可融合多个时间异步、部分观测且移动的摄像头信息。LAMP在单目基准上达到当前最佳性能,并在针对的第一人称设置中显著优于基线方法。
原文摘要 · Abstract (English)
Tracking 3D human motion from egocentric multi-camera headset is challenged by severe egomotion, partial visibility or occlusions and lack of training data. Existing methods designed for monocular video often require static or slowly-moving cameras and cannot efficiently leverage multi-view, calibrated and localized input. This makes them brittle and prone to fail on dynamic egocentric captures. We propose LAMP (Localization Aware Multi-camera People Tracking): a novel, simple framework to solve this via early disentanglement of observer and target motion. LAMP introduces a two-step process. First, we leverage the known device 6 DoF motion and calibration to convert detected 2D body keypoints from all cameras over a temporal window into a unified 3D world reference frame. Second, an end-to-end-trained spatio-temporal transformer fits 3D human motion directly to this 3D ray cloud. This "lift-then-fit" approach allows LAMP to learn and leverage a natural human motion prior in the world-space, as well as providing an elegant framework to flexibly incorporate information from multiple temporally asynchronous, partially observing and moving cameras. LAMP achieves state-of-the-art results on monocular benchmarks, while significantly outperforming baselines for our targeted egocentric setting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。