arXiv:2503.07507cs.CV2025-03中稿 · CVPR被引 9

无需微调,快速实现跨场景3D语义重建

PE3R: Perception-Efficient 3D Reconstruction

  • 前向流水线融合多视角几何与2D语义先验
  • 推理速度提升9倍,语义与几何精度达新高
  • 适合需要快速部署的3D场景理解应用

近年来,2D到3D感知的进步使得从无姿态图像中恢复3D场景语义成为可能。然而,现有方法常面临泛化能力有限、依赖场景级优化及视角间语义不一致等问题。为此,我们提出PE3R——一种无需调参的高效且可泛化的3D语义重建框架。通过在前向流水线中融合多视图几何与2D语义先验,PE3R实现了零样本跨多样场景和物体类别的泛化,无需任何场景特定微调。在开放词汇分割与多视角深度估计上的大量评估表明,PE3R不仅推理速度最快达9倍提升,还在语义与几何指标上达到新的基准。该方法为可扩展的语言驱动3D场景理解铺平道路。代码已开源:github.com/hujiecpp/PE3R。

原文摘要 · Abstract (English)

Recent advances in 2D-to-3D perception have enabled the recovery of 3D scene semantics from unposed images. However, prevailing methods often suffer from limited generalization, reliance on per-scene optimization, and semantic inconsistencies across viewpoints. To address these limitations, we introduce PE3R, a tuning-free framework for efficient and generalizable 3D semantic reconstruction. By integrating multi-view geometry with 2D semantic priors in a feed-forward pipeline, PE3R achieves zero-shot generalization across diverse scenes and object categories without any scene-specific fine-tuning. Extensive evaluations on open-vocabulary segmentation and multi-view depth estimation show that PE3R not only achieves up to 9$\times$ faster inference but also sets new state-of-the-art accuracy in both semantic and geometric metrics. Our approach paves the way for scalable, language-driven 3D scene understanding. Code is available at github.com/hujiecpp/PE3R.

3D重建语义理解多视图几何高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。