用状态空间模型提升视频人体重建的几何与运动一致性
Towards Geometry-Aware and Motion-Guided Video Human Mesh Recovery
- 采用双扫描Mamba结构建模几何约束,直接从图像特征生成可靠3D姿态序列
- 基于运动引导网络增强时序连贯性,在遮挡和模糊下仍保持高精度
- 在3DPW等三数据集上超越现有方法,且计算效率更高
现有基于视频的人体三维网格重建方法常产生物理上不合理的结果,源于其依赖有缺陷的中间3D姿态锚点,且难以有效建模复杂的时空动态。为解决这些深层架构问题,本文提出HMRMamba,首次将结构化状态空间模型(SSMs)应用于人体重建,兼具高效性与长程建模能力。框架包含两大核心贡献:其一,几何感知提升模块采用新型双扫描Mamba结构,直接利用图像特征中的几何线索进行2D到3D姿态映射,生成稳定可靠的3D姿态序列作为锚点;其二,运动引导重建网络以此锚点为基础,显式建模时间上的运动规律,显著提升最终网格的时序一致性和鲁棒性,尤其在遮挡与运动模糊场景下表现优异。在3DPW、MPI-INF-3DHP和Human3.6M三个基准上的全面评估表明,HMRMamba达到新标杆水平,在重建精度与时间一致性上均优于现有方法,并具备更优的计算效率。
原文摘要 · Abstract (English)
Existing video-based 3D Human Mesh Recovery (HMR) methods often produce physically implausible results, stemming from their reliance on flawed intermediate 3D pose anchors and their inability to effectively model complex spatiotemporal dynamics. To overcome these deep-rooted architectural problems, we introduce HMRMamba, a new paradigm for HMR that pioneers the use of Structured State Space Models (SSMs) for their efficiency and long-range modeling prowess. Our framework is distinguished by two core contributions. First, the Geometry-Aware Lifting Module, featuring a novel dual-scan Mamba architecture, creates a robust foundation for reconstruction. It directly grounds the 2D-to-3D pose lifting process with geometric cues from image features, producing a highly reliable 3D pose sequence that serves as a stable anchor. Second, the Motion-guided Reconstruction Network leverages this anchor to explicitly process kinematic patterns over time. By injecting this crucial temporal awareness, it significantly enhances the final mesh's coherence and robustness, particularly under occlusion and motion blur. Comprehensive evaluations on 3DPW, MPI-INF-3DHP, and Human3.6M benchmarks confirm that HMRMamba sets a new state-of-the-art, outperforming existing methods in both reconstruction accuracy and temporal consistency while offering superior computational efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。