arXiv:2602.23951cs.CV2026-02

无需标定即可从任意视角重建3D人体,速度快180倍

AHAP: Reconstructing Arbitrary Humans from Arbitrary Perspectives with Geometric Priors

  • 用可学习查询和软分配实现跨视角人体身份关联
  • 多视角几何融合提升人体姿态一致性与定位精度
  • 适合无标定环境下的实时3D人体重建应用

从多视角图像中重建3D人体通常需要预先标定(如棋盘格或MVS算法),限制了其在真实场景中的可扩展性。本文提出AHAP(任意视角任意人体重建),一种无需相机标定的前馈式框架。核心在于有效融合多视角几何信息,辅助人体关联、重建与定位。通过可学习的人体查询与软分配,结合对比学习监督,实现跨视角人体身份匹配;人类头部模块融合多视角特征与场景上下文,预测SMPL模型,并由多视角重投影损失约束姿态一致性。多视角几何消除了单目方法的深度模糊,通过多视角三角化获得更精确的3D人体定位。在EgoHumans和EgoExo4D数据集上的实验表明,AHAP在世界空间人体重建与相机位姿估计上均达到竞争力表现,且比基于优化的方法快180倍。

原文摘要 · Abstract (English)

Reconstructing 3D humans from images captured at multiple perspectives typically requires pre-calibration, like using checkerboards or MVS algorithms, which limits scalability and applicability in diverse real-world scenarios. In this work, we present AHAP (Reconstructing Arbitrary Humans from Arbitrary Perspectives), a feed-forward framework for reconstructing arbitrary humans from arbitrary camera perspectives without requiring camera calibration. Our core lies in the effective fusion of multi-view geometry to assist human association, reconstruction and localization. Specifically, we use a Cross-View Identity Association module through learnable person queries and soft assignment, supervised by contrastive learning to resolve cross-view human identity association. A Human Head fuses cross-view features and scene context for SMPL prediction, guided by cross-view reprojection losses to enforce body pose consistency. Additionally, multi-view geometry eliminates the depth ambiguity inherent in monocular methods, providing more precise 3D human localization through multi-view triangulation. Experiments on EgoHumans and EgoExo4D demonstrate that AHAP achieves competitive performance on both world-space human reconstruction and camera pose estimation, while being 180$\times$ faster than optimization-based approaches.

3D人体重建多视角几何无标定实时重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。