arXiv:2603.06168cs.CV2026-03被引 2

用语言查询实现全景与点云的联合语义分割,突破固定标签限制。

JOPP-3D: Joint Open Vocabulary Semantic Segmentation on Point Clouds and Panoramas

  • 融合全景图与点云数据,通过视觉-语言对齐实现开放词汇分割。
  • 在Stanford-2D-3D-s和ToF-360上优于当前最佳方法,支持任意语言查询。
  • 适合需要跨模态语义理解的机器人、自动驾驶等场景。

跨视觉模态(如3D点云与全景图像)的语义分割仍具挑战性,主要受限于标注数据稀缺及固定标签模型的适应性差。本文提出JOPP-3D框架,联合利用全景图与点云数据,实现语言驱动的场景理解。将RGB-D全景图像转换为宽视场切向视角与3D点云,通过双模态提取并对齐基础视觉-语言特征,从而支持自然语言查询生成两种输入模态的语义掩码。在Stanford-2D-3D-s与ToF-360数据集上的实验表明,该方法在开放与封闭词汇的2D和3D语义分割任务中均显著优于现有最优方法,生成结果具有一致性与语义合理性。

原文摘要 · Abstract (English)

Semantic segmentation across visual modalities such as 3D point clouds and panoramic images remains a challenging task, primarily due to the scarcity of annotated data and the limited adaptability of fixed-label models. In this paper, we present JOPP-3D, an open-vocabulary segmentation framework that jointly leverages panoramic and point cloud data to enable language-driven scene understanding. We convert RGB-D panoramic images into their corresponding wide field-of-view tangential perspectives and 3D point clouds, then use these modalities to extract and align foundational vision-language features. This allows natural language querying to generate semantic masks on both input modalities. Experimental evaluation on the Stanford-2D-3D-s and ToF-360 datasets demonstrates the capability of JOPP-3D to produce coherent and semantically meaningful segmentations across panoramic and 3D domains. Our proposed method achieves a significant improvement compared to the SOTA in open and closed vocabulary 2D and 3D semantic segmentation.

语义分割多模态开放词汇点云

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。