arXiv:2412.02903cs.CV2024-12被引 12

用第一视角视频和身体感知数据预测人体动作,提升增强现实沉浸感。

EgoCast: Forecasting Egocentric Human Pose in the Wild

  • 结合第一视角视频与身体传感器数据,实现3D姿态预测。
  • 在Ego-Exo4D挑战赛中显著超越现有方法。
  • 无需历史真实姿态,适合真实场景下的动作预测。

准确估计和预测人体姿态对提升增强现实中的用户沉浸感至关重要。本文提出EgoCast,一种基于第一视角视频和本体感知数据的双模态3D人体姿态预测方法。研究在真实动态场景下进行人体姿态预测,拓展了时序预测的边界,并在此前野外姿态估计框架基础上进一步发展。我们引入当前帧估计模块,在推理时生成伪真值姿态,从而避免传统方法在预测中依赖过去的真实姿态。在最新的Ego-Exo4D和Aria Digital Twin数据集上的实验验证了EgoCast在真实场景运动估计中的有效性。在Ego-Exo4D Body Pose 2024挑战赛中,该方法显著优于当前最先进方法,为基于第一视角输入的非剧本化活动中的人体姿态估计与预测研究奠定了基础。

原文摘要 · Abstract (English)

Accurately estimating and forecasting human body pose is important for enhancing the user's sense of immersion in Augmented Reality. Addressing this need, our paper introduces EgoCast, a bimodal method for 3D human pose forecasting using egocentric videos and proprioceptive data. We study the task of human pose forecasting in a realistic setting, extending the boundaries of temporal forecasting in dynamic scenes and building on the current framework for current pose estimation in the wild. We introduce a current-frame estimation module that generates pseudo-groundtruth poses for inference, eliminating the need for past groundtruth poses typically required by current methods during forecasting. Our experimental results on the recent Ego-Exo4D and Aria Digital Twin datasets validate EgoCast for real-life motion estimation. On the Ego-Exo4D Body Pose 2024 Challenge, our method significantly outperforms the state-of-the-art approaches, laying the groundwork for future research in human pose estimation and forecasting in unscripted activities with egocentric inputs.

姿态预测第一视角增强现实多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。