用真实数据训练,让头戴相机捕捉更准更顺滑的全身动作。
FRAME: Floor-aligned Representation for Avatar Motion from Egocentric Video
- 融合设备姿态与摄像头画面,几何对齐实现多模态融合。
- 在300帧/秒下运行,实测表现优于现有方法。
- 适合做虚拟现实、增强现实中的精准动作捕捉。
以头戴式体面双目摄像头进行第一人称动作捕捉对虚拟现实和增强现实应用至关重要,但面临严重遮挡和真实世界标注数据稀缺的挑战。现有方法依赖合成数据预训练,在真实场景中难以生成平滑准确的动作预测,尤其对下肢效果差。本文提出轻量级基于VR的数据采集方案,支持机载实时6自由度姿态追踪,并构建了迄今规模最大、动作多样性最丰富的第一人称视角真实数据集。针对设备姿态与视频流特性差异大、融合困难的问题,提出FRAME架构,通过几何合理的方式融合多模态输入,实现业界领先的身体姿态预测,可在现代硬件上达到300 FPS。最后,设计一种新型训练策略提升模型泛化能力。定性与定量评估及广泛对比均验证了方法有效性。数据、代码与CAD设计将公开于https://vcai.mpi-inf.mpg.de/projects/FRAME/
原文摘要 · Abstract (English)
Egocentric motion capture with a head-mounted body-facing stereo camera is crucial for VR and AR applications but presents significant challenges such as heavy occlusions and limited annotated real-world data. Existing methods rely on synthetic pretraining and struggle to generate smooth and accurate predictions in real-world settings, particularly for lower limbs. Our work addresses these limitations by introducing a lightweight VR-based data collection setup with on-board, real-time 6D pose tracking. Using this setup, we collected the most extensive real-world dataset for ego-facing ego-mounted cameras to date in size and motion variability. Effectively integrating this multimodal input -- device pose and camera feeds -- is challenging due to the differing characteristics of each data source. To address this, we propose FRAME, a simple yet effective architecture that combines device pose and camera feeds for state-of-the-art body pose prediction through geometrically sound multimodal integration and can run at 300 FPS on modern hardware. Lastly, we showcase a novel training strategy to enhance the model's generalization capabilities. Our approach exploits the problem's geometric properties, yielding high-quality motion capture free from common artifacts in prior works. Qualitative and quantitative evaluations, along with extensive comparisons, demonstrate the effectiveness of our method. Data, code, and CAD designs will be available at https://vcai.mpi-inf.mpg.de/projects/FRAME/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。