arXiv:2603.19929cs.CVcs.AI2026-03中稿 · CVPR被引 12

无需标记点,也能在复杂场景中稳定追踪多人3D动作。

RAM: Recover Any 3D Human Motion in-the-Wild

  • 用自适应卡尔曼滤波+语义追踪,解决遮挡和互动干扰。
  • 在PoseTrack和3DPW上零样本追踪准确率显著领先现有方法。
  • 适合真实环境下的多人3D动作捕捉,部署轻量易用。

RAM结合运动感知语义追踪器与自适应卡尔曼滤波,在严重遮挡和动态交互下实现鲁棒的身份关联。通过引入时空先验的记忆增强时序HMR模块,提升人体动作重建的一致性与平滑性。同时,轻量级预测模块可预估未来姿态以维持重建连续性,门控融合模块自适应融合重建与预测特征,确保结果连贯性与鲁棒性。在多人群体真实场景数据集PoseTrack与3DPW上的实验表明,RAM在零样本追踪稳定性与3D精度上均显著优于现有最先进方法,为无标记的野外多人3D人体动作捕捉提供通用解决方案。

原文摘要 · Abstract (English)

RAM incorporates a motion-aware semantic tracker with adaptive Kalman filtering to achieve robust identity association under severe occlusions and dynamic interactions. A memory-augmented Temporal HMR module further enhances human motion reconstruction by injecting spatio-temporal priors for consistent and smooth motion estimation. Moreover, a lightweight Predictor module forecasts future poses to maintain reconstruction continuity, while a gated combiner adaptively fuses reconstructed and predicted features to ensure coherence and robustness. Experiments on in-the-wild multi-person benchmarks such as PoseTrack and 3DPW, demonstrate that RAM substantially outperforms previous state-of-the-art in both Zero-shot tracking stability and 3D accuracy, offering a generalizable paradigm for markerless 3D human motion capture in-the-wild.

3D人体动作无标记捕捉多人群体实时追踪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。