arXiv:2602.15989cs.CV2026-02被引 72

基于新骨骼-形状解耦表示,实现单图全身3D人体建模,泛化能力强。

SAM 3D Body: Robust Full-Body Human Mesh Recovery

  • 采用骨骼与表面形状解耦的MHR网格表示,提升建模灵活性
  • 在真实场景下实现优于现有方法的精度,支持用户提示引导推理
  • 适用于需要高鲁棒性3D人体重建的研究与工业应用

我们提出SAM 3D Body(3DB),一种可提示的单图像全身3D人体网格重建模型,性能达到当前最优,在多种真实复杂场景中表现稳定且准确。3DB能同时估计人体姿态、足部和手部动作。它是首个使用新型参数化网格表示Momentum Human Rig(MHR)的模型,该表示将骨骼结构与表面形状解耦。3DB采用编码器-解码器架构,支持2D关键点和掩码等辅助提示,实现类似SAM系列模型的用户引导推理。高质量标注通过多阶段标注流程获得,结合人工关键点标注、可微优化、多视角几何与密集关键点检测。数据引擎高效筛选并处理数据,涵盖罕见姿态与特殊成像条件。我们构建了一个按姿态与外观分类的新评估数据集,支持对模型行为的细粒度分析。实验表明,3DB在定性用户偏好测试与传统定量评估中均显著优于先前方法。3DB与MHR均已开源。

原文摘要 · Abstract (English)

We introduce SAM 3D Body (3DB), a promptable model for single-image full-body 3D human mesh recovery (HMR) that demonstrates state-of-the-art performance, with strong generalization and consistent accuracy in diverse in-the-wild conditions. 3DB estimates the human pose of the body, feet, and hands. It is the first model to use a new parametric mesh representation, Momentum Human Rig (MHR), which decouples skeletal structure and surface shape. 3DB employs an encoder-decoder architecture and supports auxiliary prompts, including 2D keypoints and masks, enabling user-guided inference similar to the SAM family of models. We derive high-quality annotations from a multi-stage annotation pipeline that uses various combinations of manual keypoint annotation, differentiable optimization, multi-view geometry, and dense keypoint detection. Our data engine efficiently selects and processes data to ensure data diversity, collecting unusual poses and rare imaging conditions. We present a new evaluation dataset organized by pose and appearance categories, enabling nuanced analysis of model behavior. Our experiments demonstrate superior generalization and substantial improvements over prior methods in both qualitative user preference studies and traditional quantitative analysis. Both 3DB and MHR are open-source.

3D人体重建可提示模型网格表示真实场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。