让多人人体3D重建更准更实用,直接输出相机坐标系下的定位结果
Multi-HMR 2: Multi-Person Camera-Centric Human Detection, Mesh Recovery and Tracking

- 基于DETR框架统一处理检测、重建与跟踪,无需手工设计模块
- 在无真实相机参数情况下实现毫米级3D定位,检测精度显著提升
- 适合机器人交互、社交场景分析等需精准空间感知的应用
当前人体网格重建(HMR)多聚焦于以骨盆为中心的恢复,忽视了相机坐标系中的度量3D定位和检测精度——这对人机交互、社会场景理解等实际应用至关重要。现有评估协议常忽略这些因素,侧重个体根节点为中心的恢复,而非相机空间感知。因此,现有方法依赖固定相机假设或人工后处理,限制了鲁棒性与部署能力。本文提出Multi-HMR 2,一个简单而稳健的DETR-based框架,实现多人相机中心的人体检测、网格重建与跟踪。该模型联合预测场景一致的相机参数与人体网格,实现无需真值内参的度量3D定位。通过从SAM2蒸馏图像记忆特征,实现无视频监督的持续身份关联。尽管结构简洁——无手工组件、无视频输入、无真值相机——Multi-HMR 2在保持顶尖骨盆中心性能的同时,大幅提升了检测准确率与度量3D定位能力。
原文摘要 · Abstract (English)
Most advances in human mesh recovery (HMR) have focused on pelvis-centered recovery, overlooking metric 3D localization and detection accuracy in the camera coordinate system - two key factors for real-world applications such as human-robot interaction and social scene understanding. Current evaluation protocols often ignore these aspects, emphasizing per-person, root-centered recovery rather than camera-space perception. As a result, existing approaches rely on fixed camera assumptions or handcrafted post-processing, limiting their robustness and practical deployment. We introduce Multi-HMR 2, a simple yet robust DETR-based framework for Multi-person Camera-centric Human detection, mesh Recovery, and tracking. Multi-HMR 2 predicts a scene-consistent camera together with human meshes, enabling metric 3D localization without ground-truth intrinsics. Moreover, by distilling image-based memory features from SAM2, Multi-HMR 2 extends to tracking, achieving consistent identity association without video supervision. Despite its conceptual simplicity - no handcrafted components, no video input, and no ground-truth cameras - Multi-HMR 2 achieves state-of-the-art pelvis-centered performance while substantially improving detection accuracy and metric 3D localization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。