轻量模型SSD-Poser实现头显下实时精准全身姿态估计
SSD-Poser: Avatar Pose Estimation with State Space Duality from Sparse Observations
- 基于状态空间对偶性设计混合编码器,适配复杂动作
- 在AMASS数据集上实现高精度与极快推理速度
- 适合AR/VR实时动作捕捉,尤其对低频信号有抗抖动能力
AR/VR应用的兴起推动了从头戴设备(HMD)进行实时全身姿态估计的需求。尽管HMD能提供头和手的联合信号,但受下半身自由度限制,重建完整姿态仍具挑战。现有方法多依赖传统神经网络与生成模型(如Transformer、扩散模型),难以兼顾高精度与快速推理。为此,本文提出轻量高效模型SSD-Poser,通过精心设计的混合编码器——状态空间注意力编码器,利用状态空间对偶性建模复杂运动姿态,实现实时真实姿态重建。此外,引入频率感知解码器,有效缓解因运动信号频率变化引起的抖动问题,显著提升动作平滑性。在AMASS数据集上的全面实验表明,SSD-Poser在精度与计算效率方面均表现优异,推理速度显著优于当前最优方法。
原文摘要 · Abstract (English)
The growing applications of AR/VR increase the demand for real-time full-body pose estimation from Head-Mounted Displays (HMDs). Although HMDs provide joint signals from the head and hands, reconstructing a full-body pose remains challenging due to the unconstrained lower body. Recent advancements often rely on conventional neural networks and generative models to improve performance in this task, such as Transformers and diffusion models. However, these approaches struggle to strike a balance between achieving precise pose reconstruction and maintaining fast inference speed. To overcome these challenges, a lightweight and efficient model, SSD-Poser, is designed for robust full-body motion estimation from sparse observations. SSD-Poser incorporates a well-designed hybrid encoder, State Space Attention Encoders, to adapt the state space duality to complex motion poses and enable real-time realistic pose reconstruction. Moreover, a Frequency-Aware Decoder is introduced to mitigate jitter caused by variable-frequency motion signals, remarkably enhancing the motion smoothness. Comprehensive experiments on the AMASS dataset demonstrate that SSD-Poser achieves exceptional accuracy and computational efficiency, showing outstanding inference efficiency compared to state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。