arXiv:2603.15603cs.CV2026-03被引 6

将3D人体网格重建速度提升10倍,实现单目图像实时处理。

Fast SAM 3D Body: Accelerating SAM 3D Body for Real-Time Full-Body Human Mesh Recovery

  • 重构推理流程,支持多区域并行特征提取
  • 加速关节级姿态转换超1万倍,保持高精度重建
  • 适合实时人形机器人控制与纯视觉策略采集

SAM 3D Body(3DB)在单目3D人体网格重建中达到当前最优精度,但每张图像的推理延迟达数秒,无法满足实时应用。本文提出Fast SAM 3D Body,一个无需训练的加速框架,通过解耦串行空间依赖并采用架构感知剪枝,实现并行化的多裁剪区域特征提取与简化Transformer解码。为获取与现有类人控制及策略学习框架兼容的关节级运动学(SMPL),我们用直接前馈映射替代迭代网格拟合,使该转换加速超过10,000倍。整体框架实现高达10.9倍的端到端提速,同时保持相当的重建保真度,甚至在LSPET等基准上超越3DB。我们通过部署于纯视觉遥操作系统验证其有效性,实现了无需可穿戴IMU的实时人形控制,并直接从单个RGB流中采集操作策略。

原文摘要 · Abstract (English)

SAM 3D Body (3DB) achieves state-of-the-art accuracy in monocular 3D human mesh recovery, yet its inference latency of several seconds per image precludes real-time application. We present Fast SAM 3D Body, a training-free acceleration framework that reformulates the 3DB inference pathway to achieve interactive rates. By decoupling serial spatial dependencies and applying architecture-aware pruning, we enable parallelized multi-crop feature extraction and streamlined transformer decoding. Moreover, to extract the joint-level kinematics (SMPL) compatible with existing humanoid control and policy learning frameworks, we replace the iterative mesh fitting with a direct feedforward mapping, accelerating this specific conversion by over 10,000x. Overall, our framework delivers up to a 10.9x end-to-end speedup while maintaining on-par reconstruction fidelity, even surpassing 3DB on benchmarks such as LSPET. We demonstrate its utility by deploying Fast SAM 3D Body in a vision-only teleoperation system that-unlike methods reliant on wearable IMUs-enables real-time humanoid control and the direct collection of manipulation policies from a single RGB stream.

3D人体重建实时推理机器人控制视觉导航

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。