arXiv:2512.05988cs.CVcs.AI2025-12被引 4

用多视角联合预测3D语义点云,提升场景重建精度与效率

VG3T: Visual Geometry Grounded Gaussian Transformer

  • 基于3D高斯表示,多视角联合预测带语义的点
  • nuScenes上mIoU提升1.7个百分点,仅需46%的点数
  • 适合需要高效高精度3D重建的研究者

从多视角图像生成连贯的3D场景表征是一项基础但极具挑战的任务。现有方法在多视角融合方面表现不佳,常导致3D表示碎片化且性能欠优。为此,我们提出VG3T,一种新型多视角前馈网络,通过3D高斯表示预测3D语义占据。不同于以往从单视角图像推断高斯的方法,我们的模型以联合多视角方式直接预测一组带语义属性的高斯点。该方法克服了逐视图处理带来的碎片化与不一致问题,提供几何与语义统一表征的新范式。我们还引入网格采样和位置精修两个关键组件,缓解像素对齐高斯初始化中的距离依赖密度偏差。在nuScenes基准上,VG3T相比先前最优方法,mIoU提升1.7个百分点,同时仅使用46%的基元数量,展现出卓越的效率与性能。

原文摘要 · Abstract (English)

Generating a coherent 3D scene representation from multi-view images is a fundamental yet challenging task. Existing methods often struggle with multi-view fusion, leading to fragmented 3D representations and sub-optimal performance. To address this, we introduce VG3T, a novel multi-view feed-forward network that predicts a 3D semantic occupancy via a 3D Gaussian representation. Unlike prior methods that infer Gaussians from single-view images, our model directly predicts a set of semantically attributed Gaussians in a joint, multi-view fashion. This novel approach overcomes the fragmentation and inconsistency inherent in view-by-view processing, offering a unified paradigm to represent both geometry and semantics. We also introduce two key components, Grid-Based Sampling and Positional Refinement, to mitigate the distance-dependent density bias common in pixel-aligned Gaussian initialization methods. Our VG3T shows a notable 1.7%p improvement in mIoU while using 46% fewer primitives than the previous state-of-the-art on the nuScenes benchmark, highlighting its superior efficiency and performance.

3D重建高斯表示多视角融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。