arXiv:2512.07806cs.CV2025-12被引 6

用多视角金字塔结构,一次前向传播重建大规模3D场景。

Multi-view Pyramid Transformer: Look Coarser to See Broader

  • 分层设计:先局部后全局,再由细节到整体,兼顾效率与精度。
  • 支持数十至上百张图像输入,重建速度远超传统方法。
  • 适合需要快速、高保真3D重建的工业级应用,如虚拟现实。

我们提出多视角金字塔变换器(MVP),一种可扩展的多视角变换器架构,能通过单次前向传播直接从数十至数百张图像中重建大型3D场景。受‘看广才能见全貌,看细才能见细节’理念启发,MVP基于两大核心设计:1)从局部视图逐步扩展至组块最终覆盖全场景的局部-全局跨视图层次结构;2)从详细空间表示逐步聚合为紧凑信息密集令牌的细粒度-粗粒度内视图层次结构。该双层次结构同时实现计算高效与表征丰富,支持大规模复杂场景的快速重建。我们在多种数据集上验证了MVP性能,结果表明,当与3D高斯点云(3D Gaussian Splatting)结合时,MVP在广泛视图配置下仍保持高效率和可扩展性,且达到当前最优的泛化重建质量。

原文摘要 · Abstract (English)

We propose Multi-view Pyramid Transformer (MVP), a scalable multi-view transformer architecture that directly reconstructs large 3D scenes from tens to hundreds of images in a single forward pass. Drawing on the idea of ``looking broader to see the whole, looking finer to see the details," MVP is built on two core design principles: 1) a local-to-global inter-view hierarchy that gradually broadens the model's perspective from local views to groups and ultimately the full scene, and 2) a fine-to-coarse intra-view hierarchy that starts from detailed spatial representations and progressively aggregates them into compact, information-dense tokens. This dual hierarchy achieves both computational efficiency and representational richness, enabling fast reconstruction of large and complex scenes. We validate MVP on diverse datasets and show that, when coupled with 3D Gaussian Splatting as the underlying 3D representation, it achieves state-of-the-art generalizable reconstruction quality while maintaining high efficiency and scalability across a wide range of view configurations.

3D重建多视角变换器高效建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。