无需微调,快速实现跨场景3D语义重建
PE3R: Perception-Efficient 3D Reconstruction
- 前向流水线融合多视角几何与2D语义先验
- 推理速度提升9倍,语义与几何精度达新高
- 适合需要快速部署的3D场景理解应用
近年来,2D到3D感知的进步使得从无姿态图像中恢复3D场景语义成为可能。然而,现有方法常面临泛化能力有限、依赖场景级优化及视角间语义不一致等问题。为此,我们提出PE3R——一种无需调参的高效且可泛化的3D语义重建框架。通过在前向流水线中融合多视图几何与2D语义先验,PE3R实现了零样本跨多样场景和物体类别的泛化,无需任何场景特定微调。在开放词汇分割与多视角深度估计上的大量评估表明,PE3R不仅推理速度最快达9倍提升,还在语义与几何指标上达到新的基准。该方法为可扩展的语言驱动3D场景理解铺平道路。代码已开源:github.com/hujiecpp/PE3R。
原文摘要 · Abstract (English)
Recent advances in 2D-to-3D perception have enabled the recovery of 3D scene semantics from unposed images. However, prevailing methods often suffer from limited generalization, reliance on per-scene optimization, and semantic inconsistencies across viewpoints. To address these limitations, we introduce PE3R, a tuning-free framework for efficient and generalizable 3D semantic reconstruction. By integrating multi-view geometry with 2D semantic priors in a feed-forward pipeline, PE3R achieves zero-shot generalization across diverse scenes and object categories without any scene-specific fine-tuning. Extensive evaluations on open-vocabulary segmentation and multi-view depth estimation show that PE3R not only achieves up to 9$\times$ faster inference but also sets new state-of-the-art accuracy in both semantic and geometric metrics. Our approach paves the way for scalable, language-driven 3D scene understanding. Code is available at github.com/hujiecpp/PE3R.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。