从单张图像恢复带尺度的人体与场景三维结构,精度领先。
MetricHMSR:Metric Human Mesh and Scene Recovery from Monocular Images
- 用相机射线图提供显式度量线索,解决单目尺度模糊问题。
- 人体混合专家网络分离局部姿态与全局位置,避免误差累积。
- 以重建的人体为几何锚点,提升人-场景三维对齐精度。
我们提出MetricHMSR,一种从单张单目图像中恢复度量级人体网格与三维场景的新框架。现有方法因单目尺度模糊和弱透视相机假设难以恢复真实尺度,且完全耦合的特征表示使局部姿态与全局位移难以解耦,常需多阶段流程引入累积误差。为此,我们提出MetricHMR,通过边界相机射线图提供显式度量线索用于人体重建,并设计人体混合专家(HumanMoE)动态路由图像特征至专用专家,实现局部姿态与全局度量位置的解耦感知。利用恢复的度量人体作为几何锚点,进一步优化单目度量深度估计,实现更精确的人-场景三维对齐。大量实验表明,该方法在人体网格重建与度量人-场景重建任务上均达到当前最优性能。
原文摘要 · Abstract (English)
We introduce MetricHMSR, a novel framework for recovering metric human meshes and 3D scenes from a single monocular image. Existing methods struggle to recover metric scale due to monocular scale ambiguity and weak-perspective camera assumptions. Moreover, their fully coupled feature representations make it difficult to disentangle local pose from global translation, often requiring multi-stage pipelines that introduce accumulated errors. To address these challenges, we propose MetricHMR (Metric Human Mesh Recovery), which incorporates a bounding camera ray map representation to provide explicit metric cues for human reconstruction,together with a Human Mixture-of-Experts (HumanMoE) that dynamically routes image features to specialized experts, enabling the disentangled perception of local human pose and global metric position. Leveraging the recovered metric human as a geometric anchor, we further refine monocular metric depth estimation to achieve more accurate 3D alignment between humans and scenes.Comprehensive experiments demonstrate that our method achieves state-of-the-art performance on both human mesh recovery and metric human-scene reconstruction. Project Page: https://Metaverse-AI-Lab-THU.github.io/MetricHMSR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。