arXiv:2605.14601cs.CV2026-05中稿 · ICME 2026

用连续高斯表示提升全景3D检测精度,解决2D到3D映射失真问题。

Towards Accurate Single Panoramic 3D Detection: A Semantic Gaussian Centric Approach

论文配图:Towards Accurate Single Panoramic 3D Detection: A Semantic Gaussian Centric Approach
图 1 · 摘自论文原文
  • 以连续语义高斯体替代离散网格,保持几何连续性
  • 在Structured3D上3D AP达41.2%,优于现有方法
  • 适合需要高精度全景3D感知的自动驾驶与机器人应用

全景图像中的三维目标检测对全面场景理解至关重要,但准确地将2D特征映射到3D仍面临重大挑战。现有方法通常将2D特征投影到离散的3D网格,破坏了几何连续性并限制了表征效率。为此,本文提出PanoGSDet,一种基于连续语义3D高斯表示的单目全景3D检测框架。该框架包含全景深度估计模块和语义高斯模块。全景深度估计模块从单目全景输入中提取等距圆柱形的语义与深度特征。语义高斯模块包含语义高斯提升模块(将球面特征投影至3D语义高斯体)、语义高斯优化模块(精炼这些语义高斯体)以及高斯引导预测头(从优化后的高斯表示生成3D边界框)。在Structured3D数据集上的大量实验表明,本方法显著优于现有方法。

原文摘要 · Abstract (English)

Three-dimensional object detection in panoramic imagery is crucial for comprehensive scene understanding, yet accurately mapping 2D features to 3D remains a significant challenge. Prevailing methods often project 2D features onto discrete 3D grids, which break geometric continuity and limit representation efficiency. To overcome this limitation, this paper proposes PanoGSDet, a monocular panoramic 3D detection framework built upon continuous semantic 3D Gaussian representations. The proposed framework comprises a panoramic depth estimation component and a semantic Gaussian component. The panoramic depth estimation component extracts the equirectangular semantic and depth features from the monocular panorama input. The semantic Gaussian component includes a semantic Gaussian lifting module that projects spherical features into 3D semantic Gaussians, a semantic Gaussian optimization module that refines these semantic Gaussians, and a Gaussian guided prediction head that generates 3D bounding boxes from optimized Gaussian representations. Extensive experiments on the Structured3D dataset demonstrate that our method significantly outperforms existing methods.

3D检测全景图像高斯表示单目感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。