提出大规模多视角多人人体建模新基准与方法,解决复杂遮挡下的3D人体重建难题。
Multiview Multi-Person Human Mesh Recovery Under Large Scenes with Occlusions

- 构建包含50视角、30人交互的大型合成数据集,模拟真实复杂场景
- 通过场景级3D特征体与骨盆定位查询,实现多人3D网格精准恢复
- 引入方向与关节密度损失,有效缓解严重遮挡带来的姿态歧义
人体网格恢复(HMR)旨在从图像中重建3D人体网格。现有大多数HMR基准和方法集中在单视角多人或多视角单人重建,受限于人物数量和场景规模。此类设置难以满足真实世界中大场景和严重人物遮挡的应用需求。为此,我们提出了一个大规模合成多视角多人人体网格恢复基准,命名为MVMP-HMR。该数据集包含15个复杂场景,最多支持50个相机视角和30名相互交互的人体,具有大空间覆盖和严重遮挡,显著提升了人体网格恢复的难度。基于此基准,我们进一步提出一种多视角多人全身人体网格恢复模型,称为MVMP-HMR模型。该模型首先将多视角特征融合为场景级3D特征体,再利用3D姿态估计网络预测的骨盆关节提取个性化查询,通过交叉注意力机制从3D特征体中提取并整合每个人的3D网格。此外,我们引入两种新型损失——方向损失和3D关节密度损失,以缓解严重遮挡下的方向与姿态歧义。实验表明,现有最先进方法在所提出的MVMP-HMR基准上表现不佳,而我们的方法在大场景严重遮挡条件下持续优于已有最先进方法。
原文摘要 · Abstract (English)
Human mesh recovery (HMR) aims to recover 3D human meshes from images. Most existing HMR benchmarks and methods focus on either multi-person reconstruction from a single view or single-person reconstruction from multiple views, where the number of subjects and the scene scale are relatively limited. Such settings are insufficient for real-world applications with large scenes and severe inter-person occlusions. To address this limitation, we introduce a large-scale synthetic benchmark for multiview multi-person HMR, termed MVMP-HMR. The proposed dataset contains 15 complex scenes with up to 50 camera views and 30 interacting persons, featuring large spatial coverage and severe occlusions, which significantly increases the difficulty of human mesh recovery. Based on this benchmark, we further propose a multiview multi-person whole-body human mesh recovery model, referred to as MVMP-HMR model. The model first fuses multiview features into a scene-level 3D feature volume, and then leverages pelvis joints predicted by a 3D pose estimation network to extract person-specific queries from the 3D feature volume. These human queries are cross-attended with the 3D feature volume and integrated to decode each person's 3D mesh. Moreover, we introduce two novel losses--the orientation loss and the 3D joint density loss--to alleviate orientation and pose ambiguities under severe occlusions. Experiments demonstrate that existing state-of-the-art HMR methods struggle on the proposed MVMP-HMR benchmark, while our method consistently outperforms prior SOTAs in large-scale scenes with severe occlusions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。