让3D重建同时懂人和场景,支持真实世界应用。
HAMSt3R: Human-Aware Multi-view Stereo 3D Reconstruction
- 用多模型融合编码器理解人体与环境几何关系
- 在EgoHumans等数据集上实现高精度人体与场景联合重建
- 全前馈设计高效易用,适合真实场景部署
从少量未标定的多视角图像中恢复场景三维几何是计算机视觉的长期挑战。尽管基于学习的方法如DUSt3R和MASt3R在静态户外场景中表现优异,但在以人为中心的复杂场景中表现不佳。本文提出HAMSt3R,作为MASt3R的扩展,可从稀疏、未标定的多视角图像中实现人体与场景的联合3D重建。首先,我们利用DUNE编码器,通过蒸馏MASt3R和先进人体网格恢复模型multi-HMR的特征,提升对场景与人体的理解能力。随后引入多个网络头,实现人物分割、通过DensePose估计密集对应关系,并预测人本为中心环境的深度图,从而生成富含人体语义信息的稠密3D点云。该方法无需复杂优化流程,为全前馈结构,高效适用于实际应用。我们在EgoHumans和EgoExo4D两个包含多样化人本场景的基准上评估模型性能,并验证其在传统多视角立体重建和多视角姿态回归任务上的泛化能力。结果表明,该方法不仅能有效重建人体,同时保持在通用3D重建任务中的高性能,弥合了3D视觉中对人体与场景理解的差距。
原文摘要 · Abstract (English)
Recovering the 3D geometry of a scene from a sparse set of uncalibrated images is a long-standing problem in computer vision. While recent learning-based approaches such as DUSt3R and MASt3R have demonstrated impressive results by directly predicting dense scene geometry, they are primarily trained on outdoor scenes with static environments and struggle to handle human-centric scenarios. In this work, we introduce HAMSt3R, an extension of MASt3R for joint human and scene 3D reconstruction from sparse, uncalibrated multi-view images. First, we exploit DUNE, a strong image encoder obtained by distilling, among others, the encoders from MASt3R and from a state-of-the-art Human Mesh Recovery (HMR) model, multi-HMR, for a better understanding of scene geometry and human bodies. Our method then incorporates additional network heads to segment people, estimate dense correspondences via DensePose, and predict depth in human-centric environments, enabling a more comprehensive 3D reconstruction. By leveraging the outputs of our different heads, HAMSt3R produces a dense point map enriched with human semantic information in 3D. Unlike existing methods that rely on complex optimization pipelines, our approach is fully feed-forward and efficient, making it suitable for real-world applications. We evaluate our model on EgoHumans and EgoExo4D, two challenging benchmarks con taining diverse human-centric scenarios. Additionally, we validate its generalization to traditional multi-view stereo and multi-view pose regression tasks. Our results demonstrate that our method can reconstruct humans effectively while preserving strong performance in general 3D reconstruction tasks, bridging the gap between human and scene understanding in 3D vision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。