新数据集挑战真实复杂人体动作捕捉,暴露现有模型短板。
A Dataset and Evaluation for Complex 4D Markerless Human Motion Capture

- 构建多视角同步影像+真实3D动作真值的复杂人体运动数据集
- 现有模型在遮挡、密集交互场景下性能显著下降
- 适合研究真实场景下无标记人体动作捕捉的团队使用
基于标记的动作捕捉系统虽精度高,但依赖专用设备和标记点,难以推广。当前基准数据集缺乏真实多人交互中的复杂动态、严重遮挡及相似衣着者快速位置交换等挑战,导致领域差距。本文提出一个新数据集与评估方案,涵盖单人与多人场景,包含频繁相互遮挡、相似装扮者快速位置交换及不同距离变化的复杂动作。数据集提供同步多视角RGB与深度序列、精确相机标定、来自Vicon系统的3D动作真值,以及对应的SMPL/SMPL-X参数,确保视觉观测与动作真值精准对齐。对前沿无标记动作捕捉模型的基准测试显示,其在真实复杂条件下的性能大幅下降,凸显现有方法局限。进一步实验表明针对性微调可提升泛化能力,验证了数据集的真实性和价值。该评估揭示了现有模型的关键缺陷,并为提升鲁棒性无标记4D人体动作捕捉提供了严格基准。
原文摘要 · Abstract (English)
Marker-based motion capture (MoCap) systems have long been the gold standard for accurate 4D human modeling, yet their reliance on specialized hardware and markers limits scalability and real-world deployment. Advancing reliable markerless 4D human motion capture requires datasets that reflect the complexity of real-world human interactions. Yet, existing benchmarks often lack realistic multi-person dynamics, severe occlusions, and challenging interaction patterns, leading to a persistent domain gap. In this work, we present a new dataset and evaluation for complex 4D markerless human motion capture. Our proposed MoCap dataset captures both single and multi-person scenarios with intricate motions, frequent inter-person occlusions, rapid position exchanges between similarly dressed subjects, and varying subject distances. It includes synchronized multi-view RGB and depth sequences, accurate camera calibration, ground-truth 3D motion capture from a Vicon system, and corresponding SMPL/SMPL-X parameters. This setup ensures precise alignment between visual observations and motion ground truth. Benchmarking state-of-the-art markerless MoCap models reveals substantial performance degradation under these realistic conditions, highlighting limitations of current approaches. We further demonstrate that targeted fine-tuning improves generalization, validating the dataset's realism and value for model development. Our evaluation exposes critical gaps in existing models and provides a rigorous foundation for advancing robust markerless 4D human motion capture.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。