用超图对比学习融合多模态信息,提升密集人群3D重建精度。
Contrastive Multi-Modal Hypergraph Reasoning for 3D Crowd Mesh Recovery

- 构建多模态超图,联合语义、几何与姿态线索建模群体关系。
- 在Panoptic和GigaCrowd上达到新最优,严重遮挡下仍能恢复完整姿态。
- 适合做人群三维重建或复杂场景动作分析的研究者参考。
多人3D重建对真实交互分析至关重要,但因严重遮挡和深度模糊而困难重重。现有方法多依赖单模态输入,缺乏几何引导,且常孤立重建个体,忽略群体上下文对消除歧义的关键作用。为此,本文提出对比多模态超图推理(CoMHR),协同语义、几何与姿态线索实现群体重建。首先,通过融合RGB特征、几何先验及遮挡感知的不完整姿态初始化鲁棒节点表示;引入骨盆深度指示器作为全局空间锚点,将视觉特征与无尺度深度排序对齐。随后构建共享拓扑超图,超越成对约束,建模高阶群体动态。为增强特征融合,设计基于超图的对比学习机制,同时提升模态内区分性与跨模态正交性。该机制使网络有效传播全局上下文,即使在严重遮挡下也能推断缺失信息。在Panoptic和GigaCrowd基准上的大量实验表明,本方法取得新最佳性能。代码与预训练模型已开源。
原文摘要 · Abstract (English)
Multi-person 3D reconstruction is pivotal for real-world interaction analysis, yet remains challenging due to severe occlusions and depth ambiguity. Current approaches typically rely on single-modality inputs, which inherently lack geometric guidance. Furthermore, these methods often reconstruct subjects in isolation, neglecting the collective group context essential for resolving ambiguities in crowded scenes. To address these limitations, we propose Contrastive Multi-modal Hypergraph Reasoning to synergize semantic, geometric, and pose cues for crowd reconstruction. We first initialize robust node representations by combining RGB features, geometric priors, and occlusion-aware incomplete poses. Additionally, we introduce a pelvis depth indicator as a global spatial anchor, aligning visual features with a metric-scale-agnostic depth ordering. Subsequently, we construct a shared-topology hypergraph that moves beyond pairwise constraints to model higher-order crowd dynamics. To improve feature fusion, we design a hypergraph-based contrastive learning scheme that jointly enhances intra-modal discriminability and enforces cross-modal orthogonality. This mechanism enables the network to propagate global context effectively, allowing it to infer missing information even under severe occlusion. Extensive experiments on the Panoptic and GigaCrowd benchmarks confirm that our method achieves new state-of-the-art performance. Code and pre-trained models are available at https://github.com/SunMH-try/CoMHR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。