arXiv:2506.13040cs.CV2025-06

无需标记即可精准捕捉多人交互动作,抗遮挡能力强。

MAMMA: Markerless & Automatic Multi-Person Motion Action Capture

  • 基于分割掩码预测稠密接触感知的2D关键点,实现人物对应关系估计。
  • 在多人交互和严重遮挡下仍保持高精度,优于现有方法。
  • 构建了含密集2D关键点标注的大规模合成数据集,适合研究使用。

我们提出MAMMA,一种无需标记的多视角视频人体动作捕捉系统,可准确恢复两人交互序列中的SMPL-X参数。传统系统依赖物理标记,需专用硬件、手动贴标和大量后期处理,成本高且耗时。现有学习方法多针对单人、依赖稀疏关键点,或在遮挡与身体交互场景中表现不佳。本工作提出一种条件于分割掩码的稠密2D接触感知表面关键点预测方法,支持人物间精细对应关系估计,即使在严重遮挡下也能有效工作。采用可学习查询的新型架构,提升定位精度。为训练网络,构建大规模合成多视角数据集,融合多种来源的人体动作,涵盖极端姿态、手部运动与紧密互动,包含丰富的身体接触与遮挡,并提供带密集2D关键点标注的SMPL-X真值。结果表明,该系统无需标记即可实现接近商用标记式系统的重建质量,且无需大量人工清理。此外,针对稠密关键点预测与无标记动作捕捉缺乏通用基准的问题,我们基于真实多视角序列设计两个评估设置。数据集已公开:https://mamma.is.tue.mpg.de。

原文摘要 · Abstract (English)

We present MAMMA, a markerless motion-capture pipeline that accurately recovers SMPL-X parameters from multi-view video of two-person interaction sequences. Traditional motion-capture systems rely on physical markers. Although they offer high accuracy, their requirements of specialized hardware, manual marker placement, and extensive post-processing make them costly and time-consuming. Recent learning-based methods attempt to overcome these limitations, but most are designed for single-person capture, rely on sparse keypoints, or struggle with occlusions and physical interactions. In this work, we introduce a method that predicts dense 2D contact-aware surface landmarks conditioned on segmentation masks, enabling person-specific correspondence estimation even under heavy occlusion. We employ a novel architecture that exploits learnable queries for each landmark. We demonstrate that our approach can handle complex person--person interaction and offers greater accuracy than existing methods. To train our network, we construct a large, synthetic multi-view dataset combining human motions from diverse sources, including extreme poses, hand motions, and close interactions. Our dataset yields high-variability synthetic sequences with rich body contact and occlusion, and includes SMPL-X ground-truth annotations with dense 2D landmarks. The result is a system capable of capturing human motion without the need for markers. Our approach offers competitive reconstruction quality compared to commercial marker-based motion-capture solutions, without the extensive manual cleanup. Finally, we address the absence of common benchmarks for dense-landmark prediction and markerless motion capture by introducing two evaluation settings built from real multi-view sequences. Our dataset is available in https://mamma.is.tue.mpg.de for research purposes.

动作捕捉多人交互无标记稠密关键点

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。