用1.9万条合成步态视频提升单目视频步态分析精度
SynthGait-19K: A Physically Grounded Synthetic Video Dataset for Gait Parameter Estimation

- 基于SMPL模型统一动作捕捉数据,生成可控视角的合成步态视频
- 在真实力台数据验证下,六项步态参数估计准确率显著提升
- 适合做步态分析、人体运动恢复和跨域泛化研究的开发者
从单目视频中准确估计临床有意义的步态参数对可扩展的移动能力评估至关重要,但现有数据集受限于规模小、视角单一和视觉多样性不足。我们提出SynthGait-19k,一个基于物理的合成视频数据集,包含19,272条行走视频,源自437名受试者的6,427个动作捕捉序列,配有SMPL运动模型和六种步态参数标注。为构建该数据集,我们开发了Gait2Vid,通过SMPL统一异构动作捕捉数据,并在可控视角与场景下合成多样化RGB行走视频。我们评估了生成视频与输入步态运动学的一致性,并以力台测量验证提取的步态事件。利用SynthGait-19k,我们基准测试了直接RGB、基于姿态、生物力学及人体网格重建方法,分析了视角、训练数据规模和合成到真实域偏移的影响。我们还引入GaitXFormer作为直接RGB参考模型。合成监督在GaitXFormer和基于姿态的架构上均有效迁移至真实视频,证明其在不同表示下的实用性。我们进一步发现空间步态参数更易受视觉域偏移影响,且改进人体网格重建并不必然提升下游步态估计性能。
原文摘要 · Abstract (English)
Accurate estimation of clinically meaningful gait parameters from monocular video is important for scalable mobility assessment, yet progress is limited by the small scale, restricted viewpoints, and limited visual diversity of existing datasets. We introduce SynthGait-19k, a physically grounded synthetic video dataset containing 19,272 walking videos derived from 6,427 MoCap sequences across 437 subjects, with paired SMPL motion and annotations for six gait parameters. To construct the dataset, we develop Gait2Vid, which unifies heterogeneous MoCap recordings through SMPL and synthesizes diverse RGB walking videos under controllable viewpoints and scene appearances. We assess the generated videos for consistency with their conditioning gait kinematics and validate extracted gait events against force-platform measurements. Using SynthGait-19K, we benchmark direct RGB, pose-based, biomechanical, and human-mesh-recovery approaches and analyze viewpoint, training-data scale, and synthetic-to-real domain shift. We also introduce GaitXFormer as a direct RGB reference model for estimating gait parameters. Synthetic supervision transfers effectively to real videos across both GaitXFormer and a pose-based architecture, demonstrating utility across different representations. We further find that spatial gait parameters are more sensitive to visual domain shift and that improved HMR reconstruction alone does not necessarily translate to improved downstream gait estimation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。