arXiv:2604.08543cs.CV2026-04被引 1

用状态机提升事件相机下头戴式3D人体姿态估计精度与稳定性

E-3DPSM: A State Machine for Event-Based Egocentric 3D Human Pose Estimation

  • 设计事件驱动的连续姿态状态机,匹配事件流动态变化
  • 在两个基准上实现最高19%的误差降低和2.7倍的时间稳定性提升
  • 适合需要高精度实时动作捕捉的虚拟现实/增强现实应用

事件相机在单目头戴式3D人体姿态估计中具有毫秒级时间分辨率、高动态范围和几乎无运动模糊等优势。现有方法虽利用了这些特性,但在许多应用场景(如沉浸式VR/AR)中仍存在3D估计精度不足的问题。这主要源于其设计未充分适配事件流的异步连续特性,导致对自遮挡和时间抖动敏感。本文重新思考该任务,提出E-3DPSM——一种面向事件流的连续姿态状态机。E-3DPSM将连续人体运动与细粒度事件动态对齐,通过演化潜在状态并预测伴随事件的3D关节位置变化,再与直接的3D姿态预测融合,实现稳定且无漂移的最终3D姿态重建。该模型在单台工作站上以80 Hz实时运行,在两个基准测试中达到新最佳性能,平均关键点位置误差(MPJPE)最高降低19%,时间稳定性提升达2.7倍。

原文摘要 · Abstract (English)

Event cameras offer multiple advantages in monocular egocentric 3D human pose estimation from head-mounted devices, such as millisecond temporal resolution, high dynamic range, and negligible motion blur. Existing methods effectively leverage these properties, but suffer from low 3D estimation accuracy, insufficient in many applications (e.g., immersive VR/AR). This is due to the design not being fully tailored towards event streams (e.g., their asynchronous and continuous nature), leading to high sensitivity to self-occlusions and temporal jitter in the estimates. This paper rethinks the setting and introduces E-3DPSM, an event-driven continuous pose state machine for event-based egocentric 3D human pose estimation. E-3DPSM aligns continuous human motion with fine-grained event dynamics; it evolves latent states and predicts continuous changes in 3D joint positions associated with observed events, which are fused with direct 3D human pose predictions, leading to stable and drift-free final 3D pose reconstructions. E-3DPSM runs in real-time at 80 Hz on a single workstation and sets a new state of the art in experiments on two benchmarks, improving accuracy by up to 19% (MPJPE) and temporal stability by up to 2.7x. See our project page for the source code and trained models.

3D姿态估计事件相机状态机VR/AR

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。